One clock, three problems
For months, the BinderDex pricing dashboard had a single rule: if any card in the whole catalog was older than 30 hours, the global status flipped red. I built that rule when the catalog was small and every card refreshed on the same cadence. It felt clean: one number, one threshold, one breach. What could go wrong?
The catalog grew, and portfolio cards (the high-traffic ones people actually browse) started going stale at 12 hours when a source feed lagged. Long-tail catalog cards, which refresh on a slower schedule, routinely hit 30 hours without anything being wrong. One clock treated both as the same event. The dashboard cried wolf constantly, and somewhere inside that noise, a real emergency was waiting to get buried.
Here's what I noticed after a few weeks of this: I'd stopped reading the red. My eyes would flick to the dashboard, see the breach, and move on. Not because I didn't care, but because the signal had been diluted past the point of meaning anything. A red light that's always red is just a red light.
The split
This week I retired the single 30-hour metric and replaced it with three lane-specific freshness rules, each with its own SLA:
- Portfolio lane: 12h. These are the cards a browsing user sees. Staleness here is a user-facing breach.
- Value lane: 24h. Mid-priority cards that drift more but still need to be reasonably current.
- Catalog lane: 7d. The long tail. Stale counts stay diagnostic only, not a breach.
The breach function now asks two questions instead of one: are there unmapped portfolio cards (a card exists but has no pricing lane assigned), and are there stale cards within an active lane? The old whole-catalog count sticks around as a diagnostic number on the dashboard, but it no longer drives a breach decision. And the six per-game lane stale-card series feed the workflow dashboard directly, so I can see which game's pricing is actually lagging instead of just "something is stale somewhere."
The part I didn't expect
The change I'm happiest about wasn't in the plan. Before this, a card with no lane assignment just didn't count toward or against anything. It fell through a gap in the schema. If a card was active in the portfolio but had no lane, the old system was blind to it. Now, unmapped portfolio cards are an explicit global breach, a hard red. If a card is active in the portfolio, it has to have a lane, full stop.
I didn't find that gap by designing. I found it by splitting the clock and then asking what each lane was missing. The decomposition surfaced the unmapped-card problem on its own, because once you have lanes, a card with no lane is obviously wrong. Before, it was just invisible.
That's the thing about decomposing a metric: you almost always find a second problem hiding inside the first one. The unmapped cards were there all along. They just didn't have a shape until I gave the system lanes to fit them into.
Why this keeps happening
The pattern is bigger than freshness clocks. Any monitoring system that lumps populations with different urgency curves into one threshold will eventually do this. You start with one metric because it's simple and you're small. Then the population grows and segments, the segments age at different rates, and the threshold that was once a reasonable average becomes a lie. Not a wrong number, a number that used to mean something and now means nothing.
I think the instinct most people have when a dashboard cries wolf is to tune the threshold. Raise it from 30 hours to 48, or add a grace period. But tuning the threshold doesn't fix the problem, it just moves it. The real fix is admitting the populations are different and giving each one its own clock.
This same week, the Statpro pipeline had a callback race where a crawl framework could fire both success and failure handlers for the same request. Different problem, but the same family of failure. A jq filter passed wrong evidence because all(.; cond) checked a wrapper instead of iterating items. In all three cases, the system reported green when it shouldn't have. The alarm wasn't wrong, it was silent. And a silent alarm is the most dangerous kind, because you don't know you're not hearing it.
What I'd take away
There's a longer version of this too. The per-lane split only works if you know what your lanes are, and that means knowing which cards matter to a browsing user and which ones don't. That's a product judgment, and the dashboard can tell you something is stale. It can't tell you whether staleness matters. You have to decide that, and then wire the alarm to respect the decision.
If you have a single global threshold covering populations with different urgency, split it before it teaches you to ignore it. You won't regret the extra dashboard rows. You'll regret the fire you missed while you were reading noise.