L1halo
03 ▸ MCP server ▸ Python

Local research MCP

Claude Code was spending most of its usage re-reading the same game engine code. This MCP server hands those questions to a local model on one GPU, returns answers with checked citations, and routes the questions it gets wrong back to Claude.

Role
Design, build, eval
Stack
Python · MCP · Ollama
Model
Qwen3.8 27B IQ4_XS
Hardware
RTX 5080 16 GB
Eval score70%

on 47 engine questions. Claude Sonnet 5.5 scored 71% on the same set.

Hallucination13%

down from 26% after fixing the search tools. Sonnet: 2%.

Median answer37.5 s

down from 89.4 s after fitting all 66 layers on the GPU.

Per finished task9.6M

weighted Claude tokens, from about 20M before. Early, 3 days of data.

ClaudeBoss: a painted office where the Claude boss hands orders to local LLM workers in the basement
Screenshot pendingassets/claudeboss.jpg
ClaudeBoss, the live visualizer: Claude is the boss, local models work in the basement, subagents sit in the VIP lounge.

The problem

I mod Project Zomboid with Claude Code. Over 15 active days the mod workspace used 306.3M weighted tokens and hit the session limit 7 times. Subagents were 52% of it, mostly research workflows reading the engine's Lua and Java to answer questions like "what does this vanilla action check before it runs?"

Those questions are high-volume and read-heavy. A local model can read files for free. The open question was whether it answers correctly often enough to be worth asking.

How it works

Four MCP tools, one per job
Instant · no LLMsearch_reference

BM25 over notes the local model wrote once per file, plus hand-written engine notes.

15–60 slocal_research

Agent loop on the local model with list, search and read tools. Every citation says whether the model actually read that file.

About 60 slocal_review

First-pass review of a git diff with the same tools. Claude confirms each finding.

ClaudeRouting

High-confidence answers are used after a spot check. Medium goes back to Claude. Java engine questions go straight to Claude.

The eval

47 questions drafted from verified engine notes, with every key fact checked against the source. Each answer was graded blind by two Claude Opus judges against a reference answer, with a third judge for disagreements. Claude models ran under the same limits: Grep, Glob and Read on the same folder, at most 8 tool calls.

Local modelClaude
ModelScoreFully correctHallucinationMedian time
gemma4 12B29%4%32%11.7 s
gemma4 26B32%2%26%58.3 s
gpt-oss 20B33%9%36%6.0 s
qwen3.5 9B34%4%51%10.6 s
Claude Haiku 4.538%9%40%—
qwen3.8 27B63%34%26%89.4 s
qwen3.8 27B, fixed tools + full GPU70%45%13%37.5 s
Claude Sonnet 5.571%45%2%—

What moved the number

The tools, not the model. I replayed all 598 tool calls from one eval run. 21% of searches hit the 40-match cap and showed the model only the first 40 lines in folder order. 24% of searches with a path filter matched no file at all, and the model got "No matches", which it could not tell apart from "this function does not exist".

The fix: a filter that matches nothing now searches everything and names folders with similar names, and a capped search lists the 10 files with the most matches and puts definition lines first. Same model, same questions: 63% → 70%, hallucinations 26% → 13%.

Then the GPU. The model only fit 83–93% on the GPU because a large batch size reserved a 1.5 GiB compute buffer. Dropping it put all 66 layers on the card: generation went from 12.6 to 44 tokens/sec, with no change in score.

Where it fails

Does it generalize

On a second codebase, the ClaudeBoss visualizer in Node.js, the same model scored 95% with no hallucinations on 10 questions. That repo is about 2,000 lines with no hidden engine layer, so the gap says more about the PZ codebase than about the tools.

Real usage

The first feature built with the hybrid workflow finished 12 tasks at about 9.6M weighted tokens each, against about 20M per task in the baseline, with fewer session-limit hits per active day (0.47 → 0.33). It is 3 days of data and 4 of those tasks were one-line doc changes, so I read it as a direction, not a result yet.