The bugs you can't see until you make them prove themselves
Last week I spent ten hours on a single database migration. It adds two columns, a foreign key, and some search indexes. On paper that is one ALTER TABLE. In practice it took twelve commits, and each one fixed a different way the migration could fail without telling anyone it had failed.
Here's what surprised me: the bugs weren't in the migration. They were in the evidence system I'd built to prove the migration worked. And that system kept finding problems I would never have caught otherwise. The receipts weren't causing the bugs, they were making visible what had always been invisible.
What a receipt actually is
When I say receipt, I don't mean a log line that scrolls by. I mean a structured document, written to an immutable archive, that records exactly what was checked, what was found, and what was repaired. If a migration step runs, it must produce a receipt. If it can't produce one, the pipeline halts.
This sounds bureaucratic until you watch it catch a real bug. The migration runs inside an ECS task. My first receipt gate checked whether the task reached STOPPED with a zero exit code, which seemed straightforward. Except ECS has eight lifecycle states, and my gate only knew three. A task can be DELETED before the poller sees it, and my gate had no branch for that. It just hung, waiting for a status that would never arrive.
Without the receipt gate, the migration would have appeared to succeed. The task ran, the columns got added, life went on. But the gate refused to record a receipt for a task it couldn't confirm had stopped cleanly, and that refusal blocked the pipeline until I taught it about the full lifecycle. The receipt system was stricter than reality, and reality had to catch up.
The absent-object problem
The second bug was sneakier. The migration writes receipts to an archive bucket, and the task's IAM role couldn't tell the difference between an object that doesn't exist and an object it isn't allowed to read, because both returned access-denied. So the receipt checker would see "denied" and assume the receipt was missing, when really the receipt was right there but the role couldn't see it.
Think about what happens without this gate. The migration writes a receipt, and a later step checks for it, gets denied, and treats that as "not written." The migration looks like it failed, so you retry it. Now you have two partial receipts and a confused operator at 2am. The fix was granting list permission so the code could distinguish absence from denial explicitly. Small change, but the bug it exposed had been latent since the first deployment.
The reserved word that killed a DO block
The third one is my favorite. The migration SQL includes a precondition block that validates all ten search indexes before the main ALTER runs. Inside that block, a column alias in a SELECT INTO query collided with a Postgres reserved word. The migration would fail with a syntax error deep inside a DO block, which means the error message pointed at the wrong line and said nothing useful.
Here's the thing: this precondition existed because I wanted receipts for index validation. If I'd just run the ALTER and hoped, the indexes would have been created fine. The precondition was extra rigor I added, and the extra rigor is what surfaced the reserved-word collision. The bug was always there in the DO block, I just would never have run it independently.
Receipts as a forcing function
I started using this pattern for historical data certification. Every step of cleaning a sports season produces a receipt: how many entries checked, how many repaired, what the repairs were. When I sealed the NBA 1979-80 season last week, I could point at 1,487 checked entries and 12 repairs. Not vibes, receipts.
What I didn't expect was how the receipts would change my relationship with the code itself. When you know every step must produce verifiable evidence, you start writing code differently. You build smaller steps because each one needs its own receipt. You handle edge cases because the receipt gate will catch the ones you miss. You stop trusting logs and start trusting structured proof.
The cost is real, because migration 0024 took twelve commits and a full day. The ECS lifecycle fix, the archive permissions, the reserved word, the receipt naming collisions, the deterministic file separators. Each one was a place where my evidence system was stricter than the system it was checking. And each time, I had to decide: relax the gate, or fix the underlying code?
Almost always, the answer was fix the code. Because the gate was telling me something true: there was a state I wasn't handling, a permission I was missing, a word I was shadowing. The receipt system wasn't being pedantic, it was being honest about gaps I'd papered over.
Where this goes wrong
You can overdo it. I caught myself writing receipt gates for steps that couldn't meaningfully fail, which just added latency and false alarms. The test is simple: if the step failed silently, would downstream data corrupt? If yes, it needs a receipt. If the answer is no, a log line is fine. The migration's index preconditions need receipts because a broken index means wrong query results later. The Lambda's warmup step doesn't, because a cold Lambda just starts slower.
The other failure mode is receipts that lie. A receipt that says "checked" but didn't actually verify anything is worse than no receipt at all, because it creates false confidence. The NBA certification receipts include the actual repair records. When the AFL schedule inclusion failed a gate, it showed up as a specific failed check with a specific missing season, rather than a generic error. The receipt has to carry enough detail that you could reproduce the verification by hand.
The pattern, stripped down
If you want to try this, the shape is simple. Every step that modifies data or state produces a structured receipt, and the receipt goes somewhere immutable. The next step in the pipeline checks for the receipt before it runs. If the receipt is missing or incomplete, the pipeline halts. That's the whole thing.
The magic isn't in the checking, it's in the halting. Most systems check things and log the results, but very few refuse to proceed when the check fails. The refusal is what turns a log into a guard. A log tells you what happened after it happened. A receipt gate prevents it from happening in the first place.
I shipped a fail-closed lifecycle pipeline the same week. Sixteen commits, deployed to production, and it can't transmit a single profile. The gate is checked at every boundary that could touch data. Same instinct: don't trust the code to do the right thing. Make it impossible to do the wrong thing, and make it prove it didn't.
The ten hours on migration 0024 felt excessive at the time. Looking back, most of those hours were spent fixing bugs that had been hiding in my infrastructure for weeks. The migration didn't create them, the receipts found them. And next time, the pipeline will be a little more honest, because every gate I fixed stays fixed.