Choosing a completion model: the name on the tag decides the prompt
A practical guide to picking a fill-in-the-middle model for twinny, what the Automatic template does with the model's name, and what to check when suggestions come back wrong.
Picking a model for code completion is two decisions, and only one of them is about the model. The other is the prompt format: every family marks the text before the cursor, the text after it and the hole between them with its own tokens, and a model given another family’s tokens does something, just not completion. twinny makes that second decision from the model’s name. This guide covers how to make the first one, and how to tell when the second went wrong.
Base, not instruct
Only models trained with fill-in-the-middle (FIM) tokens can complete between a prefix and a suffix. For most families that is the variant tagged base or code. Instruct variants will produce output, but they tend to explain or chat instead of completing.
The families the docs list as tried and working:
| Family | Tags | Template |
|---|---|---|
| Qwen2.5-Coder | qwen2.5-coder:1.5b-base, :7b-base | codeqwen |
| CodeLlama | codellama:7b-code, :13b-code | codellama |
| DeepSeek Coder | deepseek-coder:1.3b-base, :6.7b-base | deepseek |
| Codestral | codestral | codestral |
| StarCoder2 | starcoder2:3b, :7b | starcoder |
| Granite Code | granite-code:3b-base, :8b-base | starcoder |
| CodeGemma | codegemma:2b-code, :7b-code | codegemma |
| Stable Code | stable-code:3b-code | stable-code |
| CodeGeeX | codegeex4 | starcoder |
On size, the docs’ advice is 1.5B to 7B: fast enough to keep up with typing, good enough to be useful. The suggested starting point is qwen2.5-coder:1.5b-base, moving up if the hardware has room. Bigger is not a safe direction in general; the docs note that the 34B CodeLlama does not do well at FIM. A suggestion that arrives after you have typed past it is worth nothing, whatever wrote it.
Two behavioural notes from the same page. StarCoder2 and CodeGemma sometimes fail to stop, and lowering twinny.temperature and twinny.maxLines helps. Qwen2.5-Coder, CodeLlama and DeepSeek Coder are the families the troubleshooting page names as stopping reliably.
What Automatic does
The autocomplete provider has a FIM template field, and it defaults to Automatic. In src/extension/completion/fim-templates.ts that is a list of substrings checked in order against the lowercased model name:
const MODEL_NAME_HINTS: [string, FimFormat][] = [
["codellama", FIM_TEMPLATE_FORMAT.codellama],
["code-llama", FIM_TEMPLATE_FORMAT.codellama],
["deepseek", FIM_TEMPLATE_FORMAT.deepseek],
["codestral", FIM_TEMPLATE_FORMAT.codestral],
["qwen3-coder", FIM_TEMPLATE_FORMAT.qwen3Coder],
["qwen", FIM_TEMPLATE_FORMAT.codeqwen],
...
["llama", FIM_TEMPLATE_FORMAT.llama]
]
Order matters: codellama is tested before llama, and qwen3-coder before qwen. The first match wins. No match falls back to the CodeLlama format.
That fallback is the thing to know about. A fine-tune published under a name of its own, or a GGUF file loaded into llama.cpp as model.gguf, contains none of those substrings. It gets <PRE> … <SUF> … <MID>, and unless it happens to be a CodeLlama derivative it has never seen those tokens. The fix is to set the template by hand in the provider form, or to name the model after its family.
The same lookup picks the stop tokens. Each family has its own list in src/common/constants/models.ts: CodeLlama stops on <EOT> and its three markers, the Qwen list has ten entries including <|endoftext|>, <|file_sep|> and the ChatML markers. So a wrong template is wrong twice: the model is prompted in a dialect it does not speak, and its real end token is not in the list that would end the reply.
Three details in the templates
An empty suffix sends no markers. With nothing after the cursor there is no middle to fill. The comment in the source says base models continue plain text reliably, whereas markers with an empty suffix make some of them, CodeLlama in particular, stop straight away. So at the end of a file the prompt is the prefix alone.
CodeLlama gets spaces Meta’s reference does not have. The reference format is <SUF>{suffix}; twinny sends <SUF> {suffix}. Per the comment, the model is far more robust with the space when the cursor sits between quotes; without it, it ignores the hole and re-emits the file from the top. Whitespace around FIM tokens is part of the format, which is worth remembering before writing a template of your own.
Some families can be told about other files. Qwen2.5-Coder, StarCoder2, Granite, CodeGemma and CodeGeeX have tokens for naming files, so neighbouring files go in as separate named blocks (<|file_sep|>path, then the text). The others get the same files as commented-out blocks above the prefix. This only matters if twinny.fileContextEnabled is on; it is off by default because it costs latency.
The exception: Qwen3-Coder
Qwen3-Coder is released only as an instruct model, which breaks the first rule above. It does fill the hole, but only when the FIM prompt arrives as the user turn of a chat, with the system message “You are a code completion assistant.” Given the bare Qwen2.5 markers it rambles.
Since 4.2.3, names containing qwen3-coder get their own template. The chat is written out as ChatML inside the prompt, so the request still goes to a completion endpoint, and Ollama is sent raw: true so it does not apply the model’s template a second time. LiteLLM gets the chat as messages. The markers stay even at the end of a file, since a chat model does not continue plain text.
The model also answers inside a markdown code block regardless of instructions, so the opening fence is dropped and the closing one ends the completion. And a half-typed word is left out of the prompt and named in the chat instead (“The completion must begin with…”), because the model treats a fragment as a typo; an answer that ignores the instruction is discarded.
When it goes wrong
Set the Twinny output channel to Debug and read the prompt and the reply. Most problems show there.
- Visible tokens in suggestions, or text that restarts the file: the template does not match the model. Set it by hand.
- Suggestions that explain themselves: an instruct model. Pull the
baseorcodetag. - Suggestions that run on: lower
twinny.maxLinesandtwinny.temperature, or change family. - Suggestions that stop mid-statement: raise
twinny.numPredictFim(512 by default). On Ollama, a small default context can also cut them; raisenum_ctx. - A model with its own format: choose custom-template and edit
~/.twinny/templates/fim.hbs. Stop words are still chosen from the model name, so an unrecognised name gets the CodeLlama stop set and the suggestion may end attwinny.maxLinesrather than at the model’s end token.
The model table is kept current at Supported models, and the template reference is under Code completion at docs.twinny.dev.