Everyone talks about the model on top — this is the part underneath, treating a successful fetch and a correct record as two different claims.
A Node.js scraper for books.toscrape.com that checks robots.txt, caches every fetch, and validates every record against a strict schema before it counts as data.
I built a Node.js scraper (native fetch, cheerio, Zod) that checks robots.txt before issuing any request, then crawls the catalogue through a disk cache in front of every fetch, an 8-second per-request timeout, and exponential backoff with jitter on 429/5xx responses (honoring Retry-After when the server sends one, never retrying 403/404). The harder problem wasn't the retry logic — it was realizing a 200 response tells you nothing about whether your CSS selector actually landed on the right node. Cheerio doesn't throw on a missed selector, it just returns an empty string, so a page with a slightly different DOM shape produces a 'successful' fetch and silently wrong output. I made every record pass a strict Zod schema before it's allowed into books.json — anything that doesn't validate is quarantined to errors.json instead of corrupting the good data — and logged every fetch, cache hit, retry, and skip to a structured run.log, so a bad record is always traceable back to the exact request that produced it.
A warm-cache run finished in 439ms versus 96,219ms cold — a 219x speedup — while returning the exact same 60 validated records, so iterating on extraction logic never means re-hitting the target site.
Zero invalid records across 60 real extractions: every field that reaches books.json has passed a strict Zod schema, and a page that renders with the wrong DOM shape gets quarantined to errors.json instead of silently corrupting output.
Every request identifies itself with a real user-agent, respects an 8-second timeout, and retries 429/5xx up to 3 times with exponential backoff plus jitter (honoring Retry-After when present) instead of hammering the site in lockstep.