A gateway for five people: one GPU box, in the order you will actually do it
A practical guide to running twinny-server for a small team, from quickstart on loopback to invite links, the queue that replaces errors on a shared GPU, and the day someone leaves.
Five developers each running a 7B model on a laptop is five warm laptops and five different model versions. One machine with a decent GPU and a server in front of it is the usual alternative. twinny-server is that server: it sits in front of the model server you already run, gives each developer a key, and records who used what.
The documentation covers every field. This post is the shorter version: what to do, in order, and which defaults matter on the first day. Everything here is in the free plan, which covers five active keys with no licence.
Start on loopback
On the machine with the models, with Node 18 or newer and a model server already running:
npx twinny-server quickstart
Quickstart writes ./twinny.gateway.json if there is none, makes an admin key, and serves. It looks for a model server on the usual local ports, asks it which models it has, and lets you pick one each for chat, autocomplete and embeddings. --yes takes the recommendations. If the models are on another host, name it: --backend lmstudio=10.0.0.5, or a URL for anything OpenAI-compatible.
Two things to notice in the output.
The admin key is printed once. Keys are stored as SHA-256 hashes, so a lost key cannot be recovered, only revoked and replaced.
The banner says listening on http://127.0.0.1:8765. Nobody else can reach it yet. That is deliberate: the admin page is reachable wherever the gateway is, so the gateway starts somewhere only you are.
Open /admin, sign in with the key, and check that the Backends panel says the backend is answering before going further. Every later problem is easier to diagnose if you know this step worked.
Let the team reach it
The gateway does not manage certificates. The docs describe two ways to expose it, and the choice is mostly about whether your team already has a private network.
If it does, a VPN or a tailnet, set listen.host in the configuration file to the machine’s address on that network and restart. The extension accepts plain http:// gateway URLs, so inside the private network that is enough.
If it does not, leave the gateway on loopback and put a reverse proxy with HTTPS in front. The example in the docs is a three-line Caddyfile that proxies a domain to 127.0.0.1:8765. Only the proxy is exposed. Every request carries the key as a bearer header, which is the reason not to skip the HTTPS part on anything reachable beyond your own network.
Either way, confirm it from a laptop before inviting anyone: /healthz on the gateway’s address should answer in a browser.
Invite, do not paste
There are three ways to give a developer a key. In order of preference:
An invite link. On the admin page under People, type a name and make a link. Check the address shown next to it is the one developers will use and not localhost. When the developer opens the link, VS Code makes the key under that name, puts it in secret storage, and shows the team’s default models to confirm. The link opens once and expires after seven days. An unopened invite holds no seat.
A sign-in code. The developer chooses Connect to team, enters the gateway URL and requests a key. VS Code shows a short code, they read it to you, and you approve it on the admin page under the name you type. Codes last ten minutes. A code creates nothing by itself; the key is made at approval. Approve only a code someone has actually read to you.
A key you made yourself, with twinny-server keys create alice, sent over a channel you trust. It works. It is also a password travelling through a chat client.
Before sending any of these, set the team’s default models under Providers & models. Connecting then configures chat, autocomplete and embeddings for the developer in one step, and nobody needs to be told which alias to pick.
One key per person. Usage is attributed per key, so a shared key makes the usage page meaningless.
What a shared GPU does under load
This is the part people worry about, and the defaults are worth knowing because they decide what a busy afternoon feels like.
limits.maxActiveRequests caps the chat and autocomplete requests running at once. The default is 4. Embeddings are not counted, because indexing a workspace is hundreds of small requests and would starve everything else.
A request that finds every slot taken waits rather than failing. The waits differ by job: by default autocomplete waits up to 500 ms (limits.queue.fimWaitMs) and chat up to 15 seconds (limits.queue.chatWaitMs). The asymmetry is the point. A completion that arrives late is a completion for a cursor that has moved; a chat answer is worth waiting for. At most 8 requests wait at once (limits.queue.maxWaiting), oldest first, and a freed slot goes to the head of the queue. An editor that moves on while its request is waiting leaves the queue without touching a backend.
Past those limits a request is refused as rate-limited, and the message says whether the queue was full or the wait ran out.
Whether 4 is the right number depends on your hardware and nothing else. /metrics serves Prometheus text to an admin key, and twinny_queued_requests shows how often the queue is in use. If it is rarely empty, raise maxActiveRequests or add a backend. If one person’s tooling is taking all the slots, limits.perKey caps concurrent requests and requests per minute for each key; it has no default.
For the scraper, make a read-only admin key: keys create metrics --admin --read-only. It can read everything and change nothing.
The day someone leaves
twinny-server keys revoke alice
Their next request fails within a second, with no restart, and the message says the key was revoked and when. Their usage history stays under their name until retention removes it; usage is kept 30 days by default. A revoked key does not count as a seat, so the seat is free at once.
When the sixth person arrives, creating the key is refused and the message names the plan. That is the point at which a licence is needed, not before.
Two pieces of housekeeping. The gateway keeps everything in one directory, ~/.twinny/server unless you moved it: key hashes, invites, usage, the audit log. Back it up along with the configuration file. And quickstart in a terminal is not a service; the operator reference has a systemd unit for that.
What this does not give you
A gateway moves the boundary of where code goes from each laptop to one machine. It does not remove the need to trust that machine, or the backend behind it. If you configure a hosted provider as a backend, prompts go to that vendor. Usage records hold the model, the duration and token counts, not content, unless recording is switched on, which needs a licence and is shown to every developer on the connect screen.
The full setup, including Docker and the Helm chart, is under Run a gateway at docs.twinny.dev.