← essays

2026-09-20 · essay

Green is not the same as safe

infraautomationdataweb

Three failures, zero alarms

Three bad things happened to systems I run this week, and every light stayed green. A CI pipeline skipped an entire lane of checks because a flag it expected never showed up. A sports data board froze for a week because one string came in a slightly different shape than the parser knew. An import page captured empty collections while telling everyone it was done. Nothing crashed, nothing turned red, and every one of them failed in the politest way possible.

Here's the claim I want to earn over the next thousand words: the most dangerous state a system can be in isn't failure. It's silence that looks like an answer. I paid three times for that lesson in seven days, and the fix was never "add monitoring." It was making the artifacts say what's missing.

The flag that never showed up

Monday night I reworked a CI pipeline that had twelve commits of drift baked in. A lint task had started scheduling itself recursively (which is exactly as fun as it sounds). Coverage got processed twice by two lanes that didn't know about each other. Jobs ran when nothing relevant had changed. Nobody noticed because the pipeline stayed green, and that was the actual problem. I made the plan testable data. One cheap job looks at changed files and emits scope flags, and one function selects the checks from those flags. Then a verifier reads every lane's summary and hard-fails the merge if the plan anyone ran doesn't match the plan anyone wrote.

Then I added the rule that matters here: every flag must be exactly true or exactly false, and missing counts as broken.

Before that, an unset flag read as "off," which sounds fine until you trace it. A lane forgets to emit its flag, its checks silently skip, the pipeline reports green, and you merge with nothing verified. The verifier now checks that every boolean it expects actually exists. If a lane doesn't speak, the merge dies with a message naming who stayed quiet. Verify the plan before you trust the run.

FINAL OVERTIME

The second failure came from one token. Games that ended in overtime arrive from the source as FINAL OVERTIME, and the strict parser stored that status as unknown. Unknown can't be finalized, can't publish a score, can't leave the board. One missing token quietly froze a week of results, and everything downstream of it just waited.

The fix was five lines. The guard around it is the interesting part: only the exact spelling maps to final. A variant with spaces or dashes stays unknown and the test suite treats any variant mapping as a regression. That looks like pedantry, but it's the opposite. If a sloppy variant maps to final, then a mid-game LIVE banner can masquerade as a finished game. A wrong final is far worse than a delayed one.

But look at the shape of the failure. The parser had two real behaviors, proceed and refuse, and it needed a third: say-it-refused. An unknown status is a refusal, and that's fine, except nothing ever surfaced the refusal as its own fact. From the outside, "the parser rejected this" and "the game hasn't ended yet" looked identical. Both looked like waiting. When you can't tell a system that's refusing from one that's still working, you can't alert on either.

The import page that lied about done

The third one hurt my favorite product the most. Paste a link into it and your collection arrives on its own. Except for a few members, for whom totals came back as NaN and the flow kept retrying into red. The screen renders in two waves: labels and placeholder boxes first, real numbers a beat later. Every signal we trusted said done while the useful data was still in flight. The first pass captured an empty collection. Then a sanity check rejected the real one, because by then the URL had picked up extra parameters.

Garbage first, false alarm second, all from trusting the paint.

How do you debug a signal that lies? I reproduced it by throttling the CPU 4x and watching the page fail in slow motion (four minutes of watching honestly beats four hours of guessing). The fix is boring on purpose: done now means the summary values on screen are real numbers. Ambiguous input fails loudly instead of inventing a collection that doesn't exist.

Treat "visible" as a claim, never as proof. Wait for the data, not the paint.

Empty is not the same as absent

All three failures share one shape, and once I saw it I couldn't unsee it (it's been showing up in everything since). Two different things both look like "nothing" from the outside: a genuine zero with evidence behind it, and the absence of any answer at all (on a dashboard they render identically, as blank). An empty result and a missing result wear the same face, and most systems treat them identically. That's exactly how silence turns into a skipped lane, a frozen week, a false done.

flowchart LR
    A["an expected answer"] -->|"evidence in the artifact"| Z["genuinely empty"]
    A -->|"nothing says anything"| M["never answered"]
    Z --> P["proceed"]
    M --> L["fail loudly"]

(The site's media uploader was down when this was drafted, so the hand-drawn version will replace this Mermaid rendering later.)

The whole design lives in that fork: evidence buys a proceed, and no evidence buys a failure you can see.

The fix that worked in all three cases was the same move, applied at different heights: put the completeness claim inside the artifact. The feed items now carry their own coverage status and a fingerprint of what the source still owes us. The board answers for itself instead of letting absence answer for it. The CI plan is asserted by a verifier that checks the flags exist before trusting what they mean. Done on the import page is defined by the data. In every case the artifact stopped being a silent quantity and became a speaking one.

Make absence loud

If you take one rule from this week, take this one: decide what silence means, on purpose, before production decides for you. Concretely, the three versions that earned their keep here. Every expected boolean is exactly true or exactly false, and missing counts as broken. Every published item carries how complete it is, and done means the numbers landed.

This stuff feels like paperwork right up until the night silence is the bug, and then it's the cheapest test you own. Alarms never fired this week because, from where the alarms sat, everything looked fine — the failures were quieter than any threshold. Green stayed green while the checks never ran, a refusal parsed like a zero, and the paint proved nothing. The scariest sound a system can make is still no sound at all.