The catch block is cheap. It is the bookkeeping that costs you.
A feed parser on Statpro threw on the first malformed item this week and killed every good item on the page. One bad row, ten good rows, all gone. I fixed it with a per-item try/catch and felt pretty smart for about five minutes, until I realized the catch was the easy part.
What actually took the afternoon was the bookkeeping. Every quarantined item now gets a structured failure record: error code, field name, the selector that matched, the item index. Without that, swallowing the error is worse than crashing. You trade a loud failure everyone notices for a silent one nobody can debug.
Most parsers default to all-or-nothing. I think that default is rational, and I want to explain why I fought it anyway.
Why all-or-nothing is the safe default
When a parser throws, the page breaks and someone notices. When it swallows errors, the page looks fine and the bad data vanishes into a gap you will not find for weeks. The all-or-nothing instinct exists because it is the only failure mode that is self-announcing. You do not need instrumentation to know you crashed.
Partial-parse is a different contract. You are saying: some items are allowed to fail, and I will catch them, and I will know which ones, and I will have enough information to fix them without reproducing the input. Every one of those clauses is a commitment. Miss any one and you have built a silent data loss machine that is strictly worse than the crash it replaced.
That is why the per-item try/catch is three lines and the bookkeeping is the rest of the function. The catch answers "what do I do with this error right now?" The bookkeeping answers "how will I debug this at 11pm on a Tuesday when I have never seen this input before?" The second question is the one that matters.
What a quarantine record actually needs
A logged error message is not enough. "Unexpected token at position 47" tells you the parser choked, which you already knew. To debug a quarantined item later, you need four things:
- Which item failed (index in the page, so you can find it in the raw HTML).
- What field triggered it (the selector or key that was being read when it threw).
- The error type (parsing failure vs. missing field vs. type mismatch, because each fix is different).
- Enough of the raw node to reproduce without re-fetching (the source HTML does not stay live forever).
Here is the shape of the loop once that exists:
for (const [itemIndex, node] of itemNodes.entries()) {
try {
items.push(parseItem($, $(node), { ...context, itemIndex }));
} catch (error) {
const failure = emptyAnalysisFailure(error, $(node));
if (!failure) throw error; // unknown errors still throw
itemFailures.push(failure);
}
}
The important line is the one that re-throws. Unknown errors are not quarantined. If the parser hits something it was not built to classify, it crashes the page, because a crash you can debug beats a swallowed error you cannot. Quarantine is for failures you have named, not for failures you have not met.
The empty page still throws
If every item on a page fails, the parser still throws. This felt wrong at first (I just built a quarantine, why not use it?), but a page where zero items survive is not a partial success. It is a signal that the source changed its markup, or the feed is down, or something structural broke. Quarantining everything would hide that signal behind a wall of individual item failures that all share one root cause.
The rule I landed on: partial-parse is for pages where most items succeed. A page where nothing succeeds is a different kind of failure and deserves a different kind of alarm.
Verify before you decorate
The same shape showed up in a second fix this week. News cards on Statpro now render player headshots and team logos, but only after a verification step confirms the player exists in the database. If verification fails, you get undecorated text. No broken image, no link to a 404.
This is the same pattern at a different scale. Quarantine is "parse what you can, record what you cannot." Verify-before-decorate is "link what you can prove, render plain text for what you cannot." Both share the underlying discipline: fail soft, but only when you can name the failure and account for it. An unverified player is not an error to hide. It is a decoration to skip, with a record of why.
When to stay all-or-nothing
I would not reach for partial-parse on a feed I controlled end to end. If the same team writes the producer and the consumer, a malformed item is a bug you should fix at the source, not paper over in the parser. Crashing is the correct pressure: it forces the fix upstream where it belongs.
Quarantine earns its keep on feeds you do not control, which for a solo founder building on top of third-party sports data is most of them. You cannot prevent a source from shipping a malformed box score at 2am. You can decide whether one bad row takes the page down. The answer should be no, but only because you built the bookkeeping that makes the no safe.
The catch block is three lines. The structured failure log, the re-throw for unknown errors, the empty-page guard, the discipline of recording enough to debug without reproducing. That is the part that took the afternoon. And it is the part that makes the catch block worth writing at all.