← essays

2026-10-04 · essay

Three AI auditors read my novel at once. The error they caught had survived dozens of reviews.

writing-with-aiai-agentspublishing

The last error in Book 1 was a number

Book 1 of Starbound went done this weekend. 63 chapter files, 163,550 words, and a review gate that finally said accepted. The manuscript is sitting with its first human reader right now. I want to talk about the last defect the process caught, because it says something uncomfortable about how quality control actually works on long documents.

The error was a fleet count. The convoy in the book has a known size, tracked in a continuity ledger: 32 hulls at departure, 30 after one loss, 29 after another. The prose told a different story. Not a wildly different one, either. Across several chapters it tallies the ships in ways you can read as including the flagship or not. Chapter by chapter, everything was consistent. The ledger never contradicted itself and the prose never contradicted itself. They just weren't counting the same fleet.

Here's what gets me. That error had already survived dozens of reviews. Every chapter went through QA gates. Every pass checked anchors, voice, continuity. Nothing found it, for months of drafts. What finally caught it: a whole-book sequential read by an AI auditor with no memory of ever having read a single chapter before.

Why every review missed it

A chapter-level review sees local consistency. Does this chapter contradict the last one? Does anyone act out of character here? Is the voice holding? Those checks are necessary, and ours ran constantly. But they share a blind spot: they all inherit context from the same rolling session, and none of them holds the whole book in view at once. Who counts hulls across 60 files? Nobody. The information "the convoy has 29 ships" existed in the ledger. The information "the prose says 32 of them made landfall" existed three chapters later. The two facts never met inside one context window until the end.

So we changed the shape of the review instead of adding more of it. The final audit contract: freeze the manuscript's SHA-256 hashes into a manifest, then commission three fresh audits, one at a time, each with no memory of the others' notes. Continuity, prose quality, character voice. Each writes an evidence-backed report with file-and-line receipts. Then one primary reviewer adjudicates: confirmed defects get bounded repairs, and every repair changes file hashes, so changed files need fresh coverage. The audit can't quietly end.

flowchart TD
    B["Book 1: 63 files, 163550 words"] --> M["frozen manifest: SHA-256 per file"]
    M --> A1["audit 1: fresh context"]
    M --> A2["audit 2: fresh context"]
    M --> A3["audit 3: fresh context"]
    A1 --> AD["adjudication: one primary reviewer"]
    A2 --> AD
    A3 --> AD
    AD -->|"defect confirmed"| R["bounded repair"]
    R --> H["re-hash changed files, fresh coverage"]
    H -.->|"changed bytes"| M
    AD -->|"nothing else found"| G["review gate accepted at manifest fef852bb"]
The whole audit. The dashed edge is the part that matters: any repair invalidates its own evidence and re-enters the loop.

The receipts, not the vibes

The frozen manifest is doing more work than it looks like. When an auditor claims a chapter says something, anyone can check that claim against the exact bytes the auditor saw. "The book got better" isn't a receipt. "The fleet count now reconciles at manifest fef852bb" is. When we refreshed a batch of chapters later, the 52 untouched chapters kept their validated receipts and didn't need re-reading. The freeze is what lets the loop terminate.

The audits found exactly one objective error in the whole manuscript. One, in 163,550 words, at the very end of a process that had been grinding for months. I have mixed feelings about that number. On one hand, three full sequential audits and the damage was a single reconcilable count. On the other hand, every cheaper check we ran, all of them, missed that one. If we'd stopped at chapter-level QA, the book would be live with an error that survives a careful human read. I read that chapter twice this month and didn't notice.

What transfers to any long document

You don't need to be writing a novel to have this problem. Any document long enough that no single review holds all of it in mind has this problem: a spec spread over 40 pages, a codebase whose conventions live in a hundred files, a year of weekly reports. The pattern from this weekend:

  • Chapter-level review catches local breakage. Keep it. It's necessary and cheap.
  • One whole-artifact sequential read by a fresh context catches the cross-document kind. It's expensive, which is why it belongs at the end, once, rather than every week.
  • Fresh context is the load-bearing part. An auditor who remembers your earlier drafts inherits your earlier mistakes as background assumptions.
  • Freeze the artifact before auditing. Hashes turn "I think it said" into checkable receipts, and they tell you exactly what a repair invalidated.

The part I keep thinking about: the defect wasn't hiding in bad prose. It was hiding in two sources of truth that agreed with themselves. That's the kind of error that grows with document size, because the distance between the sources grows too. The fix isn't reading harder. It's making sure somebody, or something, holds the whole thing at once.

Book 2 planning starts from a cleaner slate than Book 1 ever had. And the first thing on the plan is the whole-book audit, scheduled before the QA grind instead of after.