> TheAuditor / blog
• methodology, honesty

The Benchmark Does Not Write the Rules

A benchmark proves a capability exists. It never gets to define the rule. TheAuditor writes detection logic to survive real code, and tuning a rule to chase a benchmark score is forbidden by design.

There is a fast way to top a security benchmark: read the answer key and write rules that match it. The score looks incredible and the tool is useless, because it has learned the test instead of the subject. It fires on the benchmark’s exact shapes and goes quiet the moment your code phrases the same bug a little differently. This is overfitting, and in static analysis it is the difference between a demo and a product.

We hold ourselves to a principle that exists specifically to prevent it.

A benchmark proves the substrate, it never specifies the contract

The rule we enforce, in our own words: a benchmark proves the substrate, it never specifies the contract. A benchmark is allowed to show that a class of bug is detectable at all. It is never allowed to define what our detection logic says. Those are two different jobs, and we keep them apart on purpose.

So detection logic is written in the engine’s own vocabulary, against public standards, aimed at the shape of the weakness rather than the shape of any one test case. A contract for a category is written only after that category has real instances behind it, not because a benchmark has a slot for it. And tuning a rule to chase a score is not a judgment call we make case by case. It is forbidden.

Why that makes findings generalize

A rule written to a benchmark answers one question: does this exact case appear. A rule written to the weakness answers the one that matters: is this pattern of untrusted input reaching a dangerous operation present in code the author has never seen. Only the second kind survives contact with your repository, because your repository is not the benchmark and never will be. Refusing to tune toward a score is not modesty. It is the only way the number means anything once the tool leaves the lab.

How this pairs with BenchProctor

This is the same conviction from the other side. BenchProctor grades tools against an answer key the scanner is not built around, which is what makes a score on it worth citing. The benchmark’s scenarios are part of the tool’s test set, so we know a capability exists, and by construction the tool is not tuned to that key, so there is no overfitting to inherit. Any overlap between them stays qualitative: the same kinds of scenarios, not a ratio we can inflate, and the benchmark deliberately reaches past what any single scanner covers. A benchmark you cannot tune toward is the only benchmark that keeps you honest.

Honest scope note

Refusing to overfit is a constraint that costs us headline numbers we could otherwise manufacture, and we take the trade on purpose. It does not mean every category is covered. It means the categories we claim are written to generalize, and the ones we have not stated yet are reported as gaps rather than faked. What is proven is proven against an answer key the tool is never built around.

Where this sits

Building to the weakness instead of the test is the ground-truth discipline behind Code Reality Labs, kept honest by BenchProctor. TheAuditor is in final commercial release preparation and ships when its hardening checks pass. Subscribe on the main site for launch news.

Was this useful?