twinny 4.2: pull-request reviews on the GPU you already own
The 4.2 gateway reviews and triages pull requests from GitHub, GitLab, Gitea and Bitbucket with your own models, only when nobody is typing, and keeps the review on your server.
A team GPU is busy for about as long as people are typing. The rest of the day it sits there, loaded and idle. 4.2 gives it something to do: read the team’s open pull requests and review them, on the gateway, with the models you already serve.
Everything in 4.2 is on the gateway side. The extension learns only two things: to send the open workspace’s name, for routing rules, and to search the shared context index.
One list of open pulls
Switch on the GitHub, GitLab, Gitea or Bitbucket plugin, give it owner/name and a read token, and the admin page lists every open pull request across the repositories you watch, with checks, merge state and approvals as the host reports them. The token is checked against the host before it is kept and is never shown again. GitHub can use an App instead, whose private key never leaves the server. Enterprise GitHub, self-managed GitLab, Forgejo and Codeberg work by setting the host URL.
The page works out who you are from the token and tags your part in each pull, so a waiting for me view lists the pulls by other people that still need your approval. Drafts stay hidden until you switch them on.
Reviews that wait their turn
Review now sends the description and diff to one of the gateway’s chat aliases. The answer stays on the server: a summary, issues found, suggestions and a verdict, with the model, the time and who asked. If the pull has moved on since, the review says it was for an earlier commit.
Auto-review is the reason to run this on the gateway rather than in a CI job. New and updated pulls are reviewed in the background, one at a time, newest first, never drafts, and only while no developer request is running. A developer’s completion never waits behind a robot’s review. The GPU does review work in the gaps between keystrokes and overnight.
The prompt is sized for local models: at most 24k characters of description and diff, with larger patches named but left out, and at most 1,500 tokens back. Reasoning models are asked not to think (think: false on Ollama), because a review that spends its budget thinking has nothing left for the answer. A reply that is only thinking fails and says why, instead of being posted as an empty review.
Reviews go through routing like any other request and show up in usage as plugin:github and the like. If an alias has a price set, the Usage page shows what the reviews cost.
Posting back, when you choose
A finished review can be posted to the host as a comment, a change request or an approval. The post ends with a footer naming the model, so nobody mistakes it for a colleague. Auto-post does this per repository. Leave it off and reviews stay private to the admin page until someone decides they are worth sharing.
Issues get the same treatment. Triage has a model suggest labels, a possible duplicate, a priority and a first reply. Nothing is posted until someone posts the reply or applies the labels.
The rest of 4.2
- Slack, Discord and Teams webhooks for reviews, failed checks, backups and backends going down.
- SSO sign-in with any OpenID Connect provider. Signing in again replaces the key, so it doubles as lost-key recovery.
- Shared context: repositories indexed once on the gateway, searched by every developer’s chat.
- Backups to a directory or an S3-compatible bucket, optionally encrypted, restored with
twinny-server backup restore. - Audit log with each line hash-chained to the one before, read-only admins, Prometheus metrics at
/metrics, a bounded request queue, and a Helm chart.
Plugins are a licence feature: without a licence the store lists them but none can be switched on. The full list is in the changelog, and the plugins page covers setting up each forge. To try it on your own hardware, run npx twinny-server quickstart on the machine with the models. A 30-day team trial switches the plugins on.