twinny / blog

← all posts

Where the milliseconds go: tuning completion latency

A practical guide to making twinny's suggestions arrive sooner, in the order that pays off, starting with reading the log rather than changing settings.

A suggestion that arrives after you have typed past it is worth nothing, however good it was. Completion latency is the sum of four things: the pause twinny waits for, the time it takes to build the prompt, the time the server takes to read that prompt, and the time the model takes to write the reply. Each has a different fix, and most of the fixes are not settings in twinny.

This guide goes through them in the order that pays off. The first step is to measure.

Read the log first

Open the Twinny output channel (Twinny - Show logs, or the status bar menu). At the default level every completion writes two lines. One when the request goes out:

FIM #41 → qwen2.5-coder:1.5b-base · src/app.ts:120 · prompt 4312 chars

and one when it returns, with the elapsed time, the size of the result and, if twinny cut the stream, what ended it.

Two things to know about that elapsed time. The clock starts after the debounce, so the pause before the request is not in it. And it includes gathering context, so it is slightly more than the model’s own time. If the number is small and suggestions still feel late, the pause is the problem. If the number is large, look at the prompt size on the first line, then at the server.

Set the channel’s level to Debug and you also see which context blocks went into the prompt and how big each was, and lines such as FIM served from the previous suggestion (typed through) for the keystrokes that never reached the model at all.

If you run llama.cpp, its server prints per-request timings to its own console, split into prompt evaluation and generation. That split tells you which of the next sections to read.

The model and where it runs

This is most of the answer, and no setting changes it.

  • Use a small base model. A 1.5B or 3B base model on the GPU is often a better completion experience than a 7B one, even on hardware that can run the 7B. The job is narrow; the next few lines do not need a large model. Chat can afford to wait, completion cannot.
  • Check it is on the GPU. ollama ps shows what is loaded and where. A model that spilled to system RAM or the CPU still works, only slower. At the usual 4-bit quantisation, budget roughly 0.6 GB per billion parameters plus headroom for context.
  • Watch for swapping. If chat and completion use different models and the GPU cannot hold both, each request evicts the other and the next one pays the load time. Use the same model for both, use a smaller completion model, or accept the swap. Recent Ollama versions keep more than one model loaded when memory allows (OLLAMA_MAX_LOADED_MODELS).

The cold start

Ollama unloads a model after a period of inactivity, and the next request waits for it to load. Two settings deal with this.

twinny.keepAlive (default 5m) is sent with each request to Ollama. Set it to 1h, or -1 to keep the model resident until the server stops. The cost is memory that stays occupied.

twinny.warmUpModel (on by default) loads the completion model when VS Code starts or regains focus. The warm-up is a completion request with an empty prompt and a one-token cap, which is enough to make the server load the model. It is skipped when a request went to that model in the last four minutes, just under Ollama’s default keep-alive, and it only runs against servers that load on demand: Ollama, LM Studio, llama.cpp, a paired device, or an OpenAI-compatible server on this machine. A hosted API is never called. In the 4.2.7 release notes, the first completion with CodeLlama 7B on Ollama went from 11 to 17 seconds to 0.06; your numbers will depend on the model and the disk.

If the first suggestion after a break is slow and the rest are fine, this is the section that applies. If every suggestion is slow, it is not.

The prompt

A server has to read the prompt before it writes anything, and a longer prompt takes longer to read. The first log line tells you how long yours is.

What goes in:

SourceSettingSize
Lines around the cursortwinny.contextLength (100)Capped at 12,000 characters before the cursor and 3,000 after
Recent editstwinny.recentEditsEnabled (on)Up to 6 changes, 2,500 characters
IntelliSensetwinny.lspContextEnabled (on)Up to 20 items, 2,000 characters, and a 150 ms wait limit
Neighbouring filestwinny.fileContextEnabled (off)Up to 3 files, 6,000 characters

The first thing to lower is twinny.contextLength. It counts lines on each side of the cursor, so 100 means up to 200 lines of the file. Halving it makes the prompt shorter at the cost of the model seeing less of the file; on a slow machine that is usually a good trade.

The second is file context. It is off by default, but a provider marked as repository-level gathers neighbouring files whether or not the setting is on. If the Debug log lists file blocks you did not expect, check the provider.

IntelliSense context is the only source that waits on something else, the language server, and that wait is bounded at 150 ms. It runs at the same time as the other context gathering. Turning it off saves at most that, and costs the model the names it cannot see.

The reply and the pause

Generation time scales with how much the model writes. twinny already cuts the stream at the end of the enclosing block, at a stop token, or at twinny.maxLines (40), and aborts the request when it does, so the server stops generating. Beyond that:

  • twinny.multilineCompletionsEnabled off ends every suggestion at the first line break. If you only want one-liners, this is the largest saving on the reply side.
  • twinny.numPredictFim (512) is the token budget. Lowering it caps the worst case; too low and suggestions stop mid-statement.
  • A request still running after 30 seconds is abandoned and whatever arrived is kept.

Then the pause. twinny.debounceWait (300 ms) is how long twinny waits after a keystroke before sending. Lowering it makes suggestions start sooner and sends more requests that the next keystroke cancels. On a fast local model that is fine. On a slow one it keeps the GPU busy with work nobody will see. A manual trigger, Alt+\, skips the wait entirely.

twinny.completionCacheEnabled (off) remembers suggestions for identical prompts. It helps when you return to the same spot with the file unchanged, and uses memory.

In order

  1. Read the log. Find out whether the time is in the pause, the prompt or the model.
  2. Smaller base model, on the GPU, not being swapped out.
  3. Leave warm-up on; raise keepAlive if the model is unloaded between bursts.
  4. Lower contextLength; keep file context off.
  5. Single-line mode if that is all you use.
  6. Only then, debounceWait.

The settings are documented in full under Code completion, and the hardware table is in Choosing a model server.

#completion#performance#guide