The completion you did not ask for: typing through a suggestion, the cache, and when the stream stops
What twinny does on every keystroke between the ghost text appearing and the next request, why most keystrokes never reach the model, and the twelve reasons a streamed completion ends.
VS Code asks an inline completion provider for a suggestion on nearly every keystroke. If each of those became a request, a local model would spend its day generating text nobody sees. Most of twinny’s completion code is about not calling the model, and about cutting the reply short when it does. This post reads through src/extension/completion/ in the order a keystroke passes through it.
First, cancel
The first line of provideInlineCompletionItems aborts whatever is in flight. Every call bumps a request id; a run from an earlier id is told to stop, and the inference request is aborted so the server can free the GPU rather than finish generating for a cursor position that no longer exists. The same happens when the selection changes for any reason. A request that survives to the end is still checked against its id, its cancellation token and the document version before anything is shown: if the document moved on while the model was thinking, the result is logged as dropped and discarded.
Only then does twinny ask whether it needs the model at all.
Typing through the suggestion
When ghost text is shown and you type its first character, VS Code hides it and asks again. Nothing about the answer has changed; you have simply typed part of it. getSuggestionContinuation in cache.ts handles this: twinny keeps the last suggestion it showed, together with the prefix and suffix it was made for. On the next call it checks whether the current prefix ends with the beginning of that suggestion, and whether what comes before it is still the same place in the file. If so, the untyped remainder is served immediately. The log line reads served from the previous suggestion (typed through).
“The same place” is decided by anchors, not by equality. The last 500 characters of the old prefix and the first 500 of the old suffix must still match. The prefix is a sliding window of lines, so typing a newline pushes its first line out; the prefix anchor is therefore trimmed a line at a time until it fits, down to a floor of 30 characters, below which it is too generic to trust. A suggestion made at the very start of a file has nothing to anchor to and is only reused while the cursor is still there.
The continuation is also tied to a scope: the file, the language, the provider and model, the endpoint, the FIM template, and every setting that shapes the prompt or the reply (lspContextEnabled, fileContextEnabled, multilineCompletionsEnabled, maxLines, numPredictFim, temperature). Change any of them and the old suggestion is no longer an answer to the current question. Accepting a suggestion clears it too; the next request starts fresh, unless twinny.enableSubsequentCompletions is off, in which case there is no next request until you type.
The cache, and why it is narrower than it sounds
twinny.completionCacheEnabled is off by default. When on, formatted completions go into a 50-entry LRU keyed on the exact prefix and suffix, whitespace included, because whitespace matters inside strings and in Python. The key also contains the scope above and the document version.
The document version is the part to notice. Every edit increments it, so the cache never serves a stale answer into a changed file. What it does serve is the case where nothing changed: you dismiss a suggestion, move the cursor away and come back, or trigger with Alt+\ twice at the same spot. That is useful, and cheap, but it is not a cache of “prompts like this one”. If you were hoping it would make the second function in a file faster than the first, it will not, and the memory cost is why it stays off.
The requests that never start
Two more gates sit before the debounce. Mid-word, twinny does not ask: the model would only finish the word, and the language server’s suggest widget already does that. When the widget is open, VS Code cancels the plain request and asks again with the highlighted item; twinny rebuilds the prefix and suffix as if that item had been typed, so the ghost text continues from the item rather than fighting it.
Then the 300 ms wait (twinny.debounceWait), skipped on a manual trigger, and the staleness check again. Only now is a prompt built.
Deciding where the reply ends
Before the request goes out, tree-sitter parses the file to find the node under the cursor. That parse answers one question: may this completion span more than one line? Inside a string, a comment or a template literal, no. With code after the cursor on the same line, no. On a blank line, yes. After a line ending in a block opener, yes. Otherwise, one line.
The reply then streams into CompletionStream, which judges it a line at a time; a line is never cut in half. In multiline mode a completion ends when the model:
- emits one of its family’s stop tokens;
- produces 250 characters of nothing but whitespace;
- starts with a blank line while there is code after the cursor on that line: it has misread the hole, and the completion is discarded rather than tearing the line apart;
- reaches a blank line while not inside a bracket it opened itself;
- reproduces the first non-blank line of the suffix, which means it has caught up with what is already there;
- dedents out of the block the cursor is in;
- closes a bracket that was opened before the cursor line;
- hits
twinny.maxLines(40).
The bracket count ignores brackets inside string literals and after //. It is approximate, and good enough to tell “still inside something it opened” from “closed it”. A model that is asked to fill in the middle through a chat turn, as Qwen3-Coder is since 4.2.3, tends to wrap the fill in a markdown fence and to restart the cursor line; the stream drops the opening fence, ends at the closing one, and strips the repeated word fragment, or discards a reply that ignored it.
When the stream ends early the request is aborted, again to give the GPU back. If the model is still going after 30 seconds, twinny keeps what has arrived and stops waiting. Each outcome is written to the Twinny output channel as one line: the request id, how long it took, how many characters and lines survived the formatter, and which of the reasons above ended it. That last field is the quickest diagnosis when a suggestion looks truncated: ended by dedent and ended by max lines call for different fixes.
The settings named here are documented under Code completion at docs.twinny.dev.