We let the catalog pipeline use all the cores
Binderdex's idea-catalog pipeline used to run politely: one subject at a time, one render limiter per subject, CPUs idling while the machine waited on the network. A full catalog generation took the evening. Now it takes minutes, and the interesting part isn't the speed. It's which parts got parallel and which parts deliberately didn't.
The pipeline stages: fetch concepts from the LLM, build a candidate pool, assemble pages, render, embed and score. We measured first (p50 and p95 per stage), and the receipts showed the network-bound stages could run wide while the CPU-bound stages needed one shared limiter sized to the performance cores. Raising subject concurrency without it just thrashes the machine.
Three rules came out of the work:
- Network-bound stages run wide. Concepts and pool building fan out across subjects with a bounded map that returns results in input order and isolates failures (
mapBounded). - CPU-bound stages share one limiter. Assembly and rendering take from the same pool sized to performance cores, so a batch of 8 subjects can't fight each other for the CPU.
- Receipts append through one serialized queue. Concurrent writers appending JSONL lines would interleave; a failed write rejects its caller but never stalls the queue.
The core of it is small enough to show whole:
export async function mapBounded(
items: readonly T[],
limit: number,
worker: (item: T, index: number) => Promise
): Promise[]> {
const gate = pLimit(limit);
return Promise.allSettled(items.map((item, index) => gate(() => worker(item, index))));
}
Why receipts come first
Earlier runs failed in a way we couldn't explain afterwards: a crash at hour two, no idea which subjects had finished. The fix wasn't better error handling, it was receipts. Every subject appends one JSONL line (stage, timings, warnings) through the serialized queue, so a rerun can skip finished work.
The test that matters most: determinism at 1 vs 8. Concurrency changes wall time, never output. If you parallelize a pipeline, write that test first. It's the one that keeps you honest.
We let the catalog pipeline use all the cores
Binderdex's idea-catalog pipeline used to run politely: one subject at a time, one render limiter per subject, CPUs idling while the machine waited on the network. A full catalog generation took the evening. Now it takes minutes, and the interesting part isn't the speed. It's which parts got parallel and which parts deliberately didn't.
The pipeline stages: fetch concepts from the LLM, build a candidate pool, assemble pages, render, embed and score. We measured first (p50 and p95 per stage), and the receipts showed the network-bound stages could run wide while the CPU-bound stages needed one shared limiter sized to the performance cores. Raising subject concurrency without it just thrashes the machine.
Three rules came out of the work:
- Network-bound stages run wide. Concepts and pool building fan out across subjects with a bounded map that returns results in input order and isolates failures (
mapBounded). - CPU-bound stages share one limiter. Assembly and rendering take from the same pool sized to performance cores, so a batch of 8 subjects can't fight each other for the CPU.
- Receipts append through one serialized queue. Concurrent writers appending JSONL lines would interleave; a failed write rejects its caller but never stalls the queue.
The core of it is small enough to show whole:
export async function mapBounded(
items: readonly T[],
limit: number,
worker: (item: T, index: number) => Promise
): Promise[]> {
const gate = pLimit(limit);
return Promise.allSettled(items.map((item, index) => gate(() => worker(item, index))));
}
Why receipts come first
Earlier runs failed in a way we couldn't explain afterwards: a crash at hour two, no idea which subjects had finished. The fix wasn't better error handling, it was receipts. Every subject appends one JSONL line (stage, timings, warnings) through the serialized queue, so a rerun can skip finished work.
The test that matters most: determinism at 1 vs 8. Concurrency changes wall time, never output. If you parallelize a pipeline, write that test first. It's the one that keeps you honest.