One contract, three adapters: where a provider name turns into a request
Completion, chat, inline edit, review and the workspace index all end up calling the same three methods. Underneath, one adapter speaks every local server's dialect, one drives the hosted SDKs, one carries the gateway protocol. What each one adds, removes or remembers before a request goes out. From src/extension/inference.
twinny knows nineteen provider kinds: Ollama, llama.cpp, LM Studio, a generic OpenAI-compatible server, eight hosted APIs, a gateway, a paired device. The features that use them do not know any of that. src/extension/inference/types.ts opens with the rule: a feature asks for a capability, an adapter turns it into what its server wants, and “nothing above this layer knows a route, a request body or a response shape, so a new backend is a new adapter and nothing else”. This post is about what the adapters do with that freedom.
The contract
There are three capabilities, fim, chat and embeddings, plus models() for the dropdown. Everything streams: a fim or chat call returns an AsyncIterable, and a provider that only answers whole yields once, so the completion provider, the chat (and the review tab through it), the tool loop and the inline edit all read with the same for await. A chunk carries text and, when the server counts, usage; nothing is estimated, a silent backend leaves it undefined. Chat chunks can also carry a reasoning model’s thinking, a finish reason (stop or length) and tool calls.
A feature never gets a raw adapter. registry.ts wraps whatever an adapter built in guard(), which makes every method present. Ask Anthropic for embeddings and you get an unsupported-capability error before anything is sent. Every stream goes through abortable(), which races each read against the abort signal, so a cancelled completion stops the moment you type rather than when the next chunk arrives, and tells the source to release its connection. Anything a client throws comes out as an InferenceError with one of eight kinds: provider-unavailable, model-unavailable, unsupported-capability, authentication, rate-limited, timeout, cancelled, inference-failure. errors.ts decides from the HTTP status and the message (a 404 with “model” in it is model-unavailable; ECONNREFUSED is provider-unavailable), and keeps what the server actually said on cause for the log.
The registry maps kinds to adapters: the ten OpenAI-compatible kinds (Ollama, llama.cpp, LM Studio, Oobabooga, LiteLLM, Open WebUI, DeepSeek, QVAC, the generic server, and a paired device) to the HTTP adapter, the eight hosted kinds to the hosted adapter, and the gateway to the remote one. Requests that would leave the machine are wrapped once more by the secret shield.
HTTP: one switch for the wire format
adapters/http.ts is a POST to the address you configured, with a bearer header when a key is set. Completion goes through streamJsonLines(), which reads the body line by line, strips a data: prefix if there is one and skips [DONE], so Ollama’s JSON lines and an OpenAI-style server’s server-sent events come out as the same parsed objects. A connection that has not answered in sixty seconds fails with a note that the model may still be loading.
adapters/fim-dialects.ts is, as its header says, “the only place a provider name decides a wire format”. createStreamRequestBodyFim is one switch:
- Ollama, Open WebUI, a device and the generic kind get
prompt,keep_aliveand anoptionsobject withnum_predictandtemperature. When the prompt is a chat already rendered with its template, the body also saysraw: true, so Ollama does not wrap it in the model’s template a second time. - llama.cpp and Oobabooga get
promptandmax_tokensand no model name. - LM Studio gets the same with
model. - Mistral, DeepSeek and OpenRouter get the model in the body, because their completion endpoints refuse a request without it, and at most four stop sequences, because OpenAI’s completions format rejects more.
- LiteLLM’s autocomplete preset points at
/v1/chat/completions, so the prompt goes as a singlemessagesturn.
The token cap is one setting with two spellings: Ollama’s num_predict takes -1 for no limit, so it is passed as is; every body that uses max_tokens leaves a non-positive cap out, because OpenAI-style servers reject it. Reading the reply mirrors this: getFimDataFromProvider looks for the kind’s native field first (response for Ollama, content for llama.cpp), then tries every known shape, choices[0].text, choices[0].delta.content, choices[0].message.content, so a server set to the wrong kind still works if it speaks a common dialect. Token counts are pulled from whichever of three places they land: OpenAI’s usage, Ollama’s prompt_eval_count and eval_count, llama.cpp’s tokens_evaluated and tokens_predicted.
The model list is the same idea applied to GET. models() tries every plausible route at once with Promise.allSettled, six seconds each, and takes the first by preference that returns a non-empty list: /api/tags before /v1/models for Ollama, the reverse for everyone else, Open WebUI’s own /api/models first for it. An Ollama behind a proxy that only forwards /v1/models still gets a dropdown.
Chat: one client, remembered refusals
Chat on local servers and hosted APIs goes through fluency.js, a dependency that speaks to hosted APIs through their SDKs and to anything OpenAI-compatible through a base URL. adapters/fluency.ts builds one request shape and adds per-kind details:
- Messages the webview built as text parts are flattened to strings. Mistral rejects a parts list for plain text; messages with images keep their parts.
- OpenAI-compatible servers are asked for token counts with
stream_options.include_usage. A strict server that does not know the field says so by name; the request is sent again without it, and the server’s key, kind, host, port and path, goes into a set so it is not asked again. QVAC is never asked. think: falseon a request reaches Ollama asreasoning_effort: "none", which its OpenAI-style route honours. Other servers vary, so only Ollama gets it.- On Anthropic, a conversation that uses tools gets two
cache_controlmarks: the system message, which covers the tool list before it and never changes, and the newest tool result, which covers the conversation so far. The next step reads that prefix back and marks its own newest result. Plain chats are not marked. - Streamed tool calls arrive in pieces keyed by
index, the name once and the arguments as fragments; Ollama sends each call whole, llama.cpp, vLLM and LM Studio split them. A collector joins them and hands the finished calls over on the finish reason, or at the end of the stream if the server never sent one.
The hosted adapter is the same chat plus delegation. Anthropic, Gemini, Groq, Cohere and Perplexity are chat-only. Completion for Mistral and OpenRouter, and embeddings for OpenAI, are plain HTTP endpoints, so the hosted adapter hands those jobs to the HTTP adapter at the configured address. OpenAI chat that offers tools, or already holds tool calls, goes to the Responses API instead, where its reasoning models accept tools; see agent mode.
The gateway adapter carries the same three jobs over /twinny/v1 to a twinny-server, which has its own posts: one GPU box for five people and pooling a teammate’s GPU.
Why it is shaped like this
Every quirk above is a few lines in one file, named after the server that needs them. The completion provider does not know Ollama wants raw, the chat does not know Anthropic caches on request, the tool loop does not know llama.cpp splits arguments. When a server changes, or a new one turns up, the change lands in fim-dialects.ts or fluency.ts and the features do not move. It is also why the generic OpenAI-compatible kind works as often as it does: the reader of the reply tries every shape anyway.
The routes and presets per server are on docs.twinny.dev, under “How requests are made”.