What @workspace actually does: two searches, a vote, and a cross-encoder
Reading through twinny's workspace search stage by stage, from the question to the labelled blocks in the prompt, and why a 0.08 threshold means the same thing on every question.
Type @workspace in twinny’s chat and the question is answered with code you did not attach. The docs describe the pipeline in one line: embed, search twice, merge, rerank, expand, prompt. This post walks through src/extension/embeddings/ to say what each of those words hides, and what the numbers are.
What is in the index
An index is a LanceDB table under ~/.twinny/embeddings/, one folder per workspace. Each row is a chunk of a file, and a chunk is a slice along syntax boundaries: tree-sitter parses the file and syntaxBreaks walks the tree, keeping any node that fits the maximum chunk size (1,000 characters by default) as a unit and opening up bigger nodes so their children become units. A leaf that is still too big, a long string or a data blob, falls through to a plain line splitter. Files with no grammar are cut at blank lines and markdown headings. Chunks under 100 characters are merged with a neighbour, and each chunk starts with up to 100 characters of the lines before it, so a definition split across two chunks is findable from either side.
Two things happen to a chunk before it is embedded. Its relative path goes in front of the text, so the model knows that parse() in src/dates.ts is about dates. And if the embedding model is nomic-embed-text or an E5 model, the text gets that family’s task prefix (search_document: or passage:); the question later gets the matching query prefix. The comment in embedder.ts says why: those models retrieve noticeably better with the prefixes than without.
Long chunks are embedded in 600-character windows with 80 characters of overlap, one row per window. The reason is Ollama’s all-minilm, which has a 256-token window and drops everything after it without saying so. The vector search sees every window and folds them back to one hit at the rank of the nearest; the keyword index covers window zero only, so a chunk is one keyword hit.
That keyword column is not the code. It is every identifier in the chunk, lowercased and whole, followed by its parts: fetchModelEmbedding becomes fetchmodelembedding fetch model embedding, and MAX_FILE_BYTES becomes max_file_bytes max file bytes. The words of the file path are added the same way, so “where is the auth stuff” can find auth/ when nothing inside the file says auth. LanceDB builds a stemmed BM25 index over that column, which is why “chunks” finds “chunk”.
Two searches and a vote
The question goes through the same treatment. Stop words are dropped, identifiers are kept whole and split, and the result is the keyword query. Meanwhile the whole question, with its query prefix, is embedded with the index’s model.
A follow-up is detected before any of this. “And how is it tested?” has too few content words to search with, so isFollowUp (fewer than two content words, or fewer than three plus a pronoun) prepends the previous question. The files the last answer came from are also remembered.
Four searches then run at once: the nearest 20 chunks by vector distance, the best 20 by BM25, and the same two again restricted to the focus files. Focus files are the active editor, the visible editors, the files the previous answer used and the other open documents, hottest first and cut to eight. Those two passes fetch six rows each, so the file under your cursor is always among the candidates whether or not it would have made the top 20 on its own.
The four lists are merged by reciprocal rank fusion with k = 60: each list votes for its members by rank, a chunk near the top of both the semantic and the keyword list beats one that is top of only one, and the focus passes weigh their files up. The top 12 of the fused list go on. Twelve is the speed-versus-recall dial, because each candidate now costs CPU.
If the embedding server does not answer, the vector lists are empty, the keyword lists still work, and the context line under the answer says only keywords were matched.
The cross-encoder
A vector distance is relative: 0.3 can be the best of a bad list or the worst of a good one. The threshold you set in the Embeddings tab needs an absolute meaning, and that is what the reranker is for. It is a cross-encoder shipped in the extension’s models/ folder (reranker.onnx, about 87 MB, with a SentencePiece tokenizer) and run through onnxruntime-web’s wasm build. No provider is involved.
It reads the question and a candidate together, as one sequence of at most 512 tokens with the question capped at 64 so the passage keeps its tail, and outputs a logit that a sigmoid turns into a probability. The passage is the chunk with its relative path on the first line. Candidates under the threshold (0.08 by default) are dropped; the three best of those are kept as near misses so the chat can show what almost made it.
Scoring runs in worker threads so the extension host stays responsive. The pool is one worker per four cores, capped at three, because each holds its own copy of the model (about 150 MB according to the comment in reranker.ts); a candidate list is split across them. Workers idle for five minutes are shut down; the next search reloads them. If the model cannot load, scores fall back to fused order and the context line says so.
From chunks to something readable
A chunk that matched is often the middle of a function. Before the prompt is built, every kept hit is widened: the file is re-parsed (from the editor buffer if it is open, so unsaved edits count) and enclosingRange walks down from the root to the tightest node that both contains the hit and adds lines, as long as that node fits 3,000 characters. A method grows to its class when the class is small; a hit in a tiny file grows to the file. The file’s leading run of imports, up to 1,200 characters, is added once per file, after the code, and only if a hit does not already cover it.
Hits from the same file that touch or overlap are joined into one block, up to 4,000 characters, keeping the higher score. The list is cut to the snippet count (6) and then to a 12,000-character budget, skipping anything that does not fit rather than stopping at it. Each block is labelled with its file and line range, and the template in front of them tells the model to refer to those labels and to ignore any block that does not bear on the question. When the extension is connected to a gateway with the shared context plugin, the team’s index is asked at the same time and its hits are appended.
The context line under the answer reports all of this: the query as searched, how many candidates were reranked, each hit’s probability, and the near misses. Rerank threshold and Relevant code snippets take effect on the next question; the chunk sizes need a rebuild, since the windows and the vectors were made with the old ones.
Setup, the list of what gets indexed, and the tuning table are under Workspace index at docs.twinny.dev.