The number that broke my confidence in my own code
Last Monday I sat down with a question I'd been dodging: how do you know your automated entity resolver is actually right? I have a resolver that matches incoming trading card news stories to catalog card identities. It had been shipping matches for weeks. The test suite said 599 passing. Everything looked correct. And I didn't believe any of it.
Not because I found a bug. Because I realized I'd built a system that could only confirm what it already knew. The tests checked the code against fixtures I wrote. The fixtures encoded assumptions I'd made. If my assumptions were wrong, the tests would faithfully reproduce my wrongness, 599 times, in green.
So I built a calibration exercise. Not another test. A truth-blind timed review where a fresh operator sits down, looks at 50 resolved candidates with their evidence but without the ground truth, and decides: approve or reject each one. The clock runs. If you get all 50 right in under 10 minutes, the resolver passes. If you don't, it fails.
The resolver scored 44 out of 50.
What the six wrong answers looked like
The six failures weren't crashes or parse errors. They were confident wrong matches. A news story about a promotional card got matched to a base set version. A reprint announcement got matched to the original printing. These are the kind of errors that look right if you're not paying attention, which is exactly why they're dangerous. The resolver had evidence for each one. The evidence was internally consistent. It was just pointing at the wrong card.
This is the failure mode that unit tests are structurally blind to. A unit test says "given this input, produce this output." If the expected output is wrong, the test passes as long as the code reproduces it. The test is a consistency check, not a correctness check. It tells you the code matches the spec. It doesn't tell you the spec matches reality.
Why a human in the loop is the point
I could have tried to fix this with more tests. Write 50 more fixtures, cover the edge cases, ship 649 passing tests. But that's the same trap. Every fixture I write encodes my own judgment, and my judgment is exactly what's being tested. The calibration works because the operator doesn't know what I think the right answer is. They look at the evidence fresh and decide. If they agree with the resolver, that's independent confirmation. If they disagree, that's a signal I can't get from any test I wrote myself.
The exercise runs entirely offline. No database connection, no network, no production reads. A corpus of 100 owner-attested stories with known correct identities, frozen and hashed. A catalog snapshot, also hashed and bound to the corpus. A pack generator picks 50 candidates using a coverage-first selection policy that maximizes diversity across games, matchers, and relation types. The operator sees candidate metadata and resolver evidence but never the ground truth. Their decisions, timing data, and identity bindings get downloaded as strict JSON. An independent verifier reconstructs the hidden truth from the corpus and checks every decision.
The whole thing is designed so that I can't cheat even if I want to. The truth is frozen before the pack is generated. The operator can't see it. I can't see their decisions until the verifier has already scored them. The only way to pass is to actually be right.
The gap is the deliverable
Here's what surprised me most. I built the calibration expecting it to be a gate: pass or fail, ship or block. But the 44/50 score turned out to be more useful than a clean pass would have been. The six failures gave me a map. Each one pointed at a specific class of ambiguity the resolver wasn't handling. Promotional variants. Reprints across sets. Cards with near-identical names in different games. These aren't random errors. They're a taxonomy of edge cases I hadn't articulated, surfaced through disagreement instead of introspection.
I've been thinking about this a lot since. The gap between 599 passing tests and 44/50 calibration isn't a failure of testing. Tests are good at what they do. The gap is a reminder that "does the code work" has two layers, and most of us only measure one. The first layer is consistency: does the code do what the spec says. The second layer is correctness: does the spec say what reality is. Unit tests measure the first layer. Calibration measures the second.
What I'm taking from this
The calibration exercise now runs on every resolver change. It costs about 10 minutes of a fresh operator's time, which is honestly cheaper than the hour I'd spend wondering whether my new matcher broke something I can't name. And the receipt it produces is something I can point at and say "this is what 'it works' means," instead of gesturing at a green test suite.
If you're building anything that makes automated decisions about identity, matching, or classification, I'd ask you this: what's your 44 out of 50? Not your test count. Your calibration score. The number that comes from someone who doesn't know your answer checking whether your system got it right. If you don't have that number, you don't know if your code works. You know if your code matches your expectations. Those are different things, and the difference is where the bugs live.