Tombstone
harvester
Jul 2019 · age 37
What it was
A week-long tool for moving small-business websites, law firms in both samples, off an old CMS: not by copying the HTML, but by turning each site into template plus JSON so the content could be poured into a new design. It had a deliberately odd topology. Firefox was the crawler, steered by a sidebar and an injected content script Jeff called the trojan. A local Node proxy was the database: every response passed through it and was stored byte for byte before anything interpreted it. The trojan kept asking the proxy "what next?" and was told to harvest links, go to a page, or run the capture rules, a small JSON language of selectors with scoped repetition that painted a coloured border on everything it took. A replay server then served the captured site with no origin at all. It worked for one site.
Wins, for the age
- At thirty-seven, split the job into a server that decides and a browser that obeys, so dynamically loaded assets were captured for free. It found on its own the shape Puppeteer (2017) had made the standard way to crawl JavaScript-heavy sites.
- Stamped every stored row with the version of the extractor that produced it. Bump a constant and the crawl redoes exactly the stale pages, the same idea Bazel's action caching (2015) is built on.
- Stored raw bodies first and extracted second, so the replay server proved the store was a complete copy of the site.
- Built the whole loop, extension, proxy, rule walker and replay, in about 1,200 lines and one week, with eleven commits and one "w00t".
What it taught
- Store the raw thing before interpreting it, and version the interpreter.
- Do not hard-code the customer into the server. The second site needed a source edit and never left the crawl phase.
- A state machine held across a network boundary leaks where the client changes underneath it, and the human's own TODOs said so.
- Proxying is an interception strategy with rules: strip caching and cookie headers, and downgrade redirects so the browser never caches them mid-crawl.
Genealogy
Ancestors: none on record. A standalone spike. Descendants: Opinion Panel: documents taken in raw, then extracted into structured records.
Epitaph
Here lies harvester: a browser talked into crawling itself, a proxy that remembered everything it saw, and one small firm's website reborn as a pile of JSON. It worked once, on purpose.