A review in parts: what the review tab sends, skips and cuts
The extension's review tab never pastes a diff into the chat. It parses the unified diff per file, drops lockfiles and build output, trims any one file to half a part, packs the rest into at most eight requests of 16,000 characters, and writes a verdict across them. The rules, the numbers and the one setting, from src/extension/review.
There are two ways twinny reviews code, and they are easy to conflate. The gateway’s pull-request plugins review GitHub, GitLab, Gitea and Bitbucket pulls on a server, in the gaps when nobody is typing; that was the 4.2 post. The other is the review tab in the sidebar: it reviews your working tree, your branch, or a GitHub pull request, with whichever chat provider you have selected, and streams the result into the chat. The docs page covers how to use it. This post is about what happens to the diff between pressing the button and the first streamed line.
The code is src/extension/review/: git.ts runs git, service.ts talks to GitHub and the chat, and diff.ts is pure, with no VS Code or git in it, which is why it is the one with a test file (src/test/suite/review-diff.test.ts).
Three diffs, one shape
Review working tree is git diff HEAD: staged and unstaged changes to tracked files. A brand-new file does not appear until it is staged, because that is how git diff works, not a twinny choice. In a repository with no commits yet there is no HEAD, so it falls back to git diff --cached.
Review branch is git diff base...HEAD, three dots, so you get what the branch added since it forked, not every change on the base since. Picking the base is a short list in detectBaseBranch. If git knows the remote’s default branch (refs/remotes/origin/HEAD), that wins, unless it is the branch you are on. Otherwise it tries origin/main, main, origin/master, master, origin/develop, develop, origin/development, development, in that order, taking the first that exists and is not the branch you are on. The tab suggests the result; you can type another.
A GitHub pull request is fetched with the application/vnd.github.v3.diff accept header, which makes GitHub return the raw unified diff rather than JSON. The list is pulls?state=open&per_page=50. A token from twinny.githubToken, if set, goes as a bearer header; without one, public repositories work until the unauthenticated rate limit. Those two GET requests are the only calls to GitHub the extension makes. Nothing is posted back.
All three end up as the same thing: a title and a string of unified diff.
What is thrown away
parseUnifiedDiff walks the text line by line, starting a new file at each diff --git header and noting new file mode, deleted file mode, rename from/rename to, Binary files and the +/- counts. That gives one record per file with its status, its hunks and its line counts.
Then noiseReason decides which files are not worth the model’s attention. The list is short and specific:
- binary files, by the header
- lockfiles:
package-lock.json,yarn.lock,pnpm-lock.yaml,Cargo.lock,Gemfile.lock,poetry.lock,go.sum,flake.lockand the rest .min.js/.min.css, source maps,__snapshots__/and.snap- anything under
node_modules,vendor,dist,build,out,.nextortarget - assets by extension: images, fonts, archives, media, compiled libraries
- a file with no
@@hunk at all, which is a mode change or an empty file
Skipped files are not silently gone. They are listed under the file table in the chat with the reason in brackets, up to twelve, then “and N more”. If everything is noise the review does not start; you get a message saying so instead of a request about pnpm-lock.yaml.
Nothing else is filtered. Generated code outside those directories, a 3,000-line fixture, a library vendored under lib/: all reviewed. The tab does not read .gitattributes or linguist-generated; it only has the list above.
The budget
The one setting is twinny.reviewMaxDiffChars, default 16,000, floor 2,000. That is the size of one request’s diff, in characters, not tokens: the planner does not know which tokenizer your model uses, so it counts what it can. Two other numbers derive from it in service.ts:
- a single file is cut at the smaller of 8,000 characters and half a part, so lowering the setting also tightens the per-file cap
- at most eight parts, which is not configurable
Cutting a file is done at hunk boundaries. truncateFileDiff keeps whole hunks up to the cap, then appends a line saying how many lines of that file’s diff were omitted. The first hunk is always kept, even if it alone is over the cap, because a file whose entire diff is one enormous hunk is still better reviewed than dropped. The file gets a “(trimmed)” note in the table.
Packing is first-fit in file order: files go into the current part until the next one would not fit, then a new part starts. A single file bigger than a part gets a part to itself. When the eighth part is full, the rest is listed as “not reviewed (too large for one review)”. At the defaults that is a ceiling of 128,000 characters of diff per review; past that, split the change or raise the setting for a large-context model.
The setting is the one trade-off that matters. For a model with a small context window, lower it: the diff takes less of the window and the review comes in more parts. For a large-context model, raise it: fewer parts, more files in sight together, and a more coherent review than several requests to a model that has forgotten part one.
What the model sees, and what the chat keeps
Each part is one request with one user message and no system message, rendered from the review template: the title, “this is part N of M, comment only on the files shown here” when there is more than one, the asked-for structure (summary, up to eight issues with 🔴/🟠/🟡 markers and a file and line, suggestions, a one-line verdict), and the diff in a fenced block. Reviews are run one at a time; a second button press while one is streaming is refused.
When there was more than one part, a last request renders review-summary with the part reviews pasted in and asks for at most five issues across them and one line on merge readiness, told not to repeat the parts and not to invent findings. Stopping the review mid-way skips this pass.
The conversation that is saved is not the diff. Its first message is the summary summarizeReview writes: a count of files and lines, a table of up to forty reviewed files with their status and +n −m, the skipped list and the unreviewed list. The streamed reviews follow as assistant messages. So a follow-up question in that conversation (“is the null check on line 40 needed?”) goes to the model with the table and the review text, not the hunks. Which is one reason the template asks the model to quote the relevant code: what the review quotes is all a later turn will have. If a follow-up needs the actual lines, add the file with @files or paste the hunk.
The template and the summary template are both editable under Manage twinny templates; the numbers are in diff.ts and the setting. The rest is on docs.twinny.dev.