Local research MCP
Claude Code was spending most of its usage re-reading the same game engine code. This MCP server hands those questions to a local model on one GPU, returns answers with checked citations, and routes the questions it gets wrong back to Claude.
on 47 engine questions. Claude Sonnet 5.5 scored 71% on the same set.
down from 26% after fixing the search tools. Sonnet: 2%.
down from 89.4 s after fitting all 66 layers on the GPU.
weighted Claude tokens, from about 20M before. Early, 3 days of data.
assets/claudeboss.jpgThe problem
I mod Project Zomboid with Claude Code. Over 15 active days the mod workspace used 306.3M weighted tokens and hit the session limit 7 times. Subagents were 52% of it, mostly research workflows reading the engine's Lua and Java to answer questions like "what does this vanilla action check before it runs?"
Those questions are high-volume and read-heavy. A local model can read files for free. The open question was whether it answers correctly often enough to be worth asking.
How it works
Four MCP tools, one per jobBM25 over notes the local model wrote once per file, plus hand-written engine notes.
Agent loop on the local model with list, search and read tools. Every citation says whether the model actually read that file.
First-pass review of a git diff with the same tools. Claude confirms each finding.
High-confidence answers are used after a spot check. Medium goes back to Claude. Java engine questions go straight to Claude.
- A resumable indexer sends each of 1,395 vanilla Lua files to the model once. Line numbers come from regex, never from the model.
- Cached answers are only reused after the server replays the model's searches and the SHA-256 of every output still matches.
- Every call writes a log line and a live event, which feed the ClaudeBoss visualizer.
The eval
47 questions drafted from verified engine notes, with every key fact checked against the source. Each answer was graded blind by two Claude Opus judges against a reference answer, with a third judge for disagreements. Claude models ran under the same limits: Grep, Glob and Read on the same folder, at most 8 tool calls.
| Model | Score | Fully correct | Hallucination | Median time |
|---|---|---|---|---|
| gemma4 12B | 29% | 4% | 32% | 11.7 s |
| gemma4 26B | 32% | 2% | 26% | 58.3 s |
| gpt-oss 20B | 33% | 9% | 36% | 6.0 s |
| qwen3.5 9B | 34% | 4% | 51% | 10.6 s |
| Claude Haiku 4.5 | 38% | 9% | 40% | — |
| qwen3.8 27B | 63% | 34% | 26% | 89.4 s |
| qwen3.8 27B, fixed tools + full GPU | 70% | 45% | 13% | 37.5 s |
| Claude Sonnet 5.5 | 71% | 45% | 2% | — |
What moved the number
The tools, not the model. I replayed all 598 tool calls from one eval run. 21% of searches hit the 40-match cap and showed the model only the first 40 lines in folder order. 24% of searches with a path filter matched no file at all, and the model got "No matches", which it could not tell apart from "this function does not exist".
The fix: a filter that matches nothing now searches everything and names folders with similar names, and a capped search lists the 10 files with the most matches and puts definition lines first. Same model, same questions: 63% → 70%, hallucinations 26% → 13%.
Then the GPU. The model only fit 83–93% on the GPU because a large batch size reserved a 1.5 GiB compute buffer. Dropping it put all 66 layers on the card: generation went from 12.6 to 44 tokens/sec, with no change in score.
Where it fails
- It still hallucinates about 6 times as often as Sonnet. That is why Claude spot-checks the claim its code depends on before using an answer.
- Questions whose answer lives in the Java engine score 55%. The index covers Lua only, so those route to Claude.
- In the review smoke test it found 2 of 4 planted bugs and missed a logic bug. Claude still reviews every change.
- The confidence label is useful: high-confidence answers score 82%, so the routing rule trusts those and escalates the rest.
Does it generalize
On a second codebase, the ClaudeBoss visualizer in Node.js, the same model scored 95% with no hallucinations on 10 questions. That repo is about 2,000 lines with no hidden engine layer, so the gap says more about the PZ codebase than about the tools.
Real usage
The first feature built with the hybrid workflow finished 12 tasks at about 9.6M weighted tokens each, against about 20M per task in the baseline, with fewer session-limit hits per active day (0.47 → 0.33). It is 3 days of data and 4 of those tasks were one-line doc changes, so I read it as a direction, not a result yet.
L1haloOpen ▸