twinny / blog

← all posts

twinny 4.2.7: the model never sees your keys, and the first completion no longer waits

A secret shield swaps credentials for placeholders before a prompt leaves the machine and puts them back in the reply, in the extension and now on the gateway. And the completion model is loaded before you type.

Two small releases, 4.2.6 and 4.2.7, with one idea between them: the text that leaves your machine should not contain your credentials, even when you paste them in yourself. The second release also fixes the slowest completion of the day, the first one.

What was leaking

A coding assistant sees a lot of secrets by accident. The .env file you attached to a chat to ask why the app will not start. The AWS key in the fixture next to the cursor, sent as context for a completion. The connection string with a password in it, pasted into a question about a timeout. If the model is on your own machine, none of that matters. If it is a hosted API, a gateway, a paired device or a server elsewhere on the network, it has just left the building.

twinny’s answer is not to refuse the request, and not to ask you first. It rewrites the prompt.

The shield

Before a prompt goes out, every credential in it is replaced with a placeholder: REDACTED_GITHUB_TOKEN_1, REDACTED_PRIVATE_KEY_1, and so on. The same secret gets the same placeholder everywhere in one request, so the model can still reason about “the key on line 3” and write code that uses it. When the reply comes back, the placeholders are swapped for the real values. You see working code with your key in it; the model never saw the key.

This covers chat, completions, inline edit and embeddings. Detection is deliberately conservative. It knows the shapes of GitHub, GitLab, AWS, Stripe, Slack, OpenAI, Anthropic, Google, Hugging Face and npm tokens, PEM private key blocks, JSON web tokens and passwords in URLs. Beyond those, a quoted value assigned to a name that says it is secret (api_key, password, client_secret) is withheld too, in code and in .env files, but only if it looks real: process.env.STRIPE_KEY, <your-key>, changeme and xxxxxxxx are left alone, and so is anything with too little entropy to be a key.

A chat reply says N secrets withheld and which kinds. Completions log the same line to the Twinny output channel. Nothing is stored.

The twinny.secretShield setting decides when it runs. The default, offMachine, shields anything that is not on this computer: hosted APIs, a gateway, a paired device, a server on another host. always adds local servers, for the cautious. off sends prompts as they are.

Now on the gateway too

4.2.7 puts the same shield in twinny-server. The gateway swaps credentials for placeholders before a request reaches a backend and restores them in the reply, for every client: the extension, the TUI, Neovim, anything that speaks to it. An operator sets it once, with policy.secretShield in the configuration file or on the admin page’s Policy tab, and it holds for the whole team whatever each developer has in their own settings.

The modes are the same. offMachine, the default, covers hosted APIs, backends on other hosts and the pooled teammates’ GPUs. always adds backends on the gateway’s own host. off forwards prompts unchanged. The setting stays on the gateway; the extension is not told about it and does not need to be.

The first completion of the day

A local model takes seconds to load, and the first completion after you open VS Code, or after lunch, waits for all of it. With CodeLlama 7B on Ollama the wait was 11 to 17 seconds. Now it is 0.06.

The trick is small. When VS Code starts or regains focus, twinny sends the completion model an empty prompt with a one-token cap. The server loads the model and answers at once, and by the time you type, it is in memory. Nothing is sent if the model was used in the last four minutes, just under Ollama’s default five-minute keep-alive, so a model about to be unloaded is touched again instead. A warm-up that takes more than two minutes is abandoned quietly; the next real completion reports the server as down if it is.

This is only for servers that load models on demand and keep them in memory a while: Ollama, LM Studio, llama.cpp, and an OpenAI-compatible server on this machine. A hosted API is always warm and bills every call, so it is never touched. twinny.warmUpModel turns it off.

Where to read more

Both releases are in the changelog. The extension updates itself from the Marketplace; the gateway is npx twinny-server@latest, or the ghcr.io/twinnydotdev/twinny-server image. If you run a team gateway, the operations page covers policy.

#release#security#completion#gateway