twinny / blog

← all posts

twinny 4.3: agent mode, one key away, and a chat that behaves like a terminal

Six releases in two days. The chat model can now read, search and edit files and run commands in your workspace, behind a switch rather than a setting; commands run in the background with their output in the reply; messages typed mid-reply queue instead of vanishing; and the gateway's Recordings fold an agent conversation into one row.

4.3.0 shipped on 1 October and 4.3.5 on 2 October. Between them the chat gained agent mode: the model can list, read, search and edit files, ask the language servers about the code, read git history and run commands, with every step shown in the reply where it happened. The changelog introduces this as a switch at the bottom left of the composer, which undersells it: the whole tools loop in src/extension/tools/ is new in 4.3.0. This post is about what it does, what it will not do, and the patch releases that followed.

Agent mode is marked experimental. The setting’s own description says it needs a capable coder model, such as qwen3-coder:30b, with a context of 8k tokens or more, because small models tend to call tools badly.

A switch, not a setting

Agent mode is turned on from a switch in the composer or with Shift+Tab. While it is on the prompt becomes a bright ❯❯ that pulses while the model works. The choice is kept for every window in the extension’s global storage rather than in settings, and src/webview/hooks/useAgentMode.ts says why: a settings write makes VS Code reload configuration, which the user sees as a flicker. Until you first flip the switch, the twinny.chatTools setting decides, and it is off by default. It is a machine-scoped setting read from your user settings only, so a workspace you clone cannot turn agent mode on for you.

With it on, the model gets a fixed set of tools over the first workspace folder: listing, finding and reading files, grep, search_code through the workspace index if you have built one, a read-only git, the language servers’ diagnostics, references, definitions and renames, editor_context, file edits, creates, moves and deletes, and run_command. A reply takes at most twelve steps (DEFAULT_MAX_STEPS in tools/loop.ts); then the model is asked for its answer.

Files ignored by a .gitignore, .git, node_modules and anything matching twinny.embeddingIgnoredGlobs are off limits to every tool, git and search_code included. Edits apply straight away by default and are saved, so Ctrl+Z in the file undoes one; twinny.chatToolsEdits set to review opens each as a diff instead, as an inline edit does. A few things wait for you whatever the setting: deleting a file git has no copy of, and changing .vscode/ or a .gitignore.

Tool calling uses the server’s own mechanism where it has one; when a server or model refuses, the model is asked to write its calls in the reply instead, and the output channel says so. 4.3.0 also fixed a bug this surfaced: a tool with no parameters, such as editor_context, streams no arguments, and the adapter for Anthropic, Bedrock, Gemini and Cohere parsed that empty string as JSON before the next request, which failed with “Unexpected end of JSON input”. Empty arguments are now {} in inference/adapters/fluency.ts, and tools/loop.ts does the same for arguments that will not parse, since what goes back to the server has to parse there too.

Commands

4.3.0 ran the model’s commands in a twinny tools terminal. 4.3.2 moved them to the background by default; the implementation is in src/extension/chat/tool-sinks.ts.

Each command is spawned in a process of its own from the workspace root, through /bin/bash where there is one (the lines models write assume a POSIX shell, which fish is not) and the system shell otherwise. Standard input is ignored, so a command that waits for a prompt gets end-of-file rather than hanging; PAGER and GIT_PAGER are cat and NO_COLOR is set. Output streams into the step as it prints. After two minutes (COMMAND_TIMEOUT_MS) the command is stopped and the model gets what there is. The stop is a SIGTERM to the child’s process group, or taskkill /T on Windows, so whatever the command started ends with it. The model is given the exit code and the tail of the output, at most 80 lines or 4000 characters, with ANSI stripped. Servers and watch modes therefore belong in a terminal of your own; twinny.chatToolsCommandsRunIn set to terminal brings the old behaviour back, provided VS Code’s shell integration is available.

Whether a command asks first is decided per command, in commandMode in chat/index.ts, in this order. If you have ever flipped the auto-run switch next to the agent switch, its position wins. Otherwise twinny.chatToolsCommands applies: ask (the default), allow or off, and off also hides the switch. Otherwise, a command you once answered with Always run runs unasked: that choice appends the normalised command to a list in global storage, matched exactly, and only Twinny - Forget commands set to always run clears it. Any other command still asks.

When one is waiting and the composer is empty, Enter runs it, Shift+Enter always runs it and Esc skips it; a second Esc stops the reply. A running command has a stop button on its line. It ends that command and anything it started, the model is told you stopped it and gets the output so far, and the reply carries on.

The chat as a terminal

The other half of 4.3.0 is the keyboard. Ctrl+C stops a reply, or clears the draft when nothing is streaming; with text selected it still copies. Esc twice clears the draft and ↑ brings it back. Ctrl+L starts a new conversation and PgUp and PgDn scroll the transcript. ? lists them. These are the chat’s own keys, not VS Code keybindings, so Ctrl is Ctrl on macOS too.

Pressing Enter while a reply was running used to clear what you had typed without sending it, which in agent mode meant losing a paragraph. Messages now wait under the transcript marked queued and go out one per finished reply. Stopping the reply yourself puts the queue back into the composer instead, after whatever is already there.

The patch releases

4.3.1: a file the agent read, shown with line numbers in a narrow chat, broke every word into its own column. Long lines now wrap as text.

4.3.3 is on the gateway. Agent mode sends the conversation again for every tool step, so one reply filled Recordings with a dozen near-identical rows. groupRows in src/gateway/ui/recordings.tsx now folds the records of one thread into a single row with a step count, placed where its newest step falls, since autocompletes interleave with them. Storage is unchanged: each step is still its own record, and the training export still has one example per step.

4.3.4 undoes a bit of 4.3.2. Fitting the composer footer into a narrow panel hid the provider and model dropdowns’ lists, so clicking one did nothing.

4.3.5 is housekeeping. Edit twinny templates was in the command palette but never registered, so it failed with “command not found”; it now opens ~/.twinny/templates, and a test checks that every contributed command is registered after activation. The es-CL translations were registered under the key esCL, so choosing Chilean Spanish gave you general Spanish. twinny.numPredictChat is gone: nothing had read it since 3.21. And twinny.temperature now says it applies to completions only; chat, inline edit and review send no temperature.

The agent mode page at docs.twinny.dev has the full tool table and settings; the changelog in the repository has every change per release.

#release#chat#agent-mode#gateway