Every Answer Says What It Left Out
We stress-tested TheAuditor's MCP server against our own codebase with the questions an agent asks mid-task, and tightened every answer that came back short. Each answer now comes from stored facts and says when it was cut, and a gate steers the agent to the query instead of the file.
An agent cannot tell a partial answer from a complete one unless the answer tells it. If a tool returns fifty rows and says nothing else, the agent reasons as if there are fifty. If it returns nothing, the agent concludes that nothing is there. A code-intelligence tool that trims its answers without saying so has the same flaw as a scanner that hands back a silent zero, and we have written about that one before.
So we turned the MCP server on our own codebase and asked it the questions an agent actually asks in the middle of a task: who calls this, what does this file declare, where does this value go, what depends on this module. Then we checked every answer against the code, and tightened each one that came back short of it.
What the stress test tightened
Answers are now read from the facts the index stores, not rebuilt from strings after the fact. On top of that:
- Java call questions. Callers, callees, and impact for a Java method resolve through the same identities the call graph uses, so they return everything the index holds.
- Totals that are totals. A result larger than what fits says so in plain words: “Showing 50 of 65.” A capped list never presents the cap as the count.
- Visible trimming. When rows or text are cut to fit, the answer marks it.
- Focused flow questions. A flow question is pinned to one variable inside one function, so steps from unrelated functions that happen to share a name stay out.
- Dependents counted as files. A file that imports five names from one module counts as one dependent, not five.
- Tier limits named. When a license tier withholds rows, the answer says the tier did it, so nobody chases a limit that will not help.
The rule underneath is one we already hold for findings: an incomplete answer is fine, an incomplete answer that presents itself as complete is not.
A gate instead of a reminder
The second lesson came from watching real sessions. A reminder to query instead of reading competes with an agent’s habit of opening files, and habit tends to win.
So in Claude Code, setup now installs a gate. A whole-file read of source code, a grep or glob across code, or a shell command that reads a code file is declined, and the refusal names the query to use instead. Reading the span of lines you are about to edit still works, small files can still be read whole, and documentation stays readable. Nothing is hidden from the agent; the gate sends it to the query that answers its question.
By default the agent now sees one lookup tool plus a status tool instead of twenty separate ones, which keeps the fixed description it carries on every request small. The full set of single-purpose tools is one setup flag away for anyone who wants it.
Honest scope note
This makes answers honest about their edges. It does not make every question answerable: where the index holds no fact for something, the answer says so rather than guessing. The read gate was built as a Claude Code hook first; Codex connects to the same server and tools.
Where this sits
Verified answers are what the ground-truth layer of Code Reality Labs is for: an agent that queries facts instead of guessing from fragments. TheAuditor is in final commercial release preparation and ships when its hardening checks pass. Subscribe on the main site for launch news.
Was this useful?