Self-Healing Scrapers: Let One Agent Write Another Agent's Instructions
A scraper usually breaks the moment a website redesigns its layout. This one doesn't — it researches the page itself, and re-researches it when something changes This article walks through building exactly that, with Claude Code.
The business need
Let's imagine a situation where parents have three kids in three different schools. None of the schools agree on how to publish a timetable. The formats are different: one puts up clean, per-class HTML pages. Another exports its entire grade level as a single PDF. What the parents want is one combined, printable schedule covering all three kids. Writing an HTML scraper for the first school and a PDF scraper for the second is an afternoon's work. What is more, the timetable structures may change in time.
The process
The solution I landed on with Claude Code was a self-correcting process based on LLM agents. There are two agents involved. The first, a research agent, studies a school's site and writes plain-English instructions for the fetch agent to follow. The fetch agents are created for each school separately. The fetch agent, follows those instructions to actually pull a class's schedule. There's one more rule tying them together, call it the self-correcting scraper process: if a fetch agent's site ever stops matching what it was told to expect, it calls the research agent back to work out what changed.
This post walks through how that pipeline is built. Use this article as an example on how to write self-correcting webpage scrapig process.
Here's the whole thing as a sequence diagram — who calls whom, and when the self-correction loop kicks in:
Research agent - phase 1
The research agent (schedule-research) never touches end-user data. Its entire job is to look at one target site and answer two questions: what classes does this school offer, and how, mechanically, do you get one of them. Given a starting URL, it:
- Locates the actual data page.
- Enumerates every class the site offers and records each one's own URL — never invented, only what was actually observed.
- Picks one real class and fetches it, working out the concrete extraction rules: table structure, field mapping, day/time formats, anything site-specific (merged cells, split groups, quirky abbreviations).
- Writes those rules down as a brand-new agent's system prompt —
fetch-<school-id>.md— with a fixed, load-bearing clause at the end that Phase 2 depends on.
The key design choice: the research agent writes instructions, in plain English, not code. That matters for two reasons. First, an LLM agent reading "the third column is the room code" generalizes better across a site's own inconsistencies than a hard-coded selector would. Second, it means the "regeneration" step later — when the site changes — is just rewriting a markdown file, not regenerating and testing a script.
Worth noting: I didn't hand-write the research agent's own instructions either. I asked Claude Code directly to design it — describe the two roles, how the agent produced in step 4 should call the research agent back when it breaks, and what each of them is and isn't responsible for — and let the model draft the actual schedule-research.md prompt from that description. The research agent's own prompt is itself LLM-generated.
The Self-Correcting Scraper Process
The self-correcting scraper process itself is a single clause, copied into every fetch agent:
---
name: fetch-<school-id>
description: Fetches a class's schedule from <school> and writes it as JSON.
tools: WebFetch, Read, Write, Task
---
## Self-correction
If the expected structure isn't found (selector/table missing, unexpected
layout, empty or clearly-wrong result), do not guess or fabricate data.
Invoke the `schedule-research` agent, naming this school and exactly what
you saw instead of what you expected. Once it rewrites this file, re-read
your own updated instructions and retry once. If it still fails, report
the failure clearly instead of returning guessed data.
That's it. No error-code taxonomy, no retry-with-backoff framework — just an instruction to name the failure and ask the agent that knows how to fix it. The research agent, invoked this way, re-examines just the part that broke (not a full from-scratch rediscovery) and rewrites the fetch agent's own file. The fetch agent then re-reads its own updated instructions and tries again. Two agents, one shared file, one rule about who to call.
The output of Phase 1 is a brand-new fetch agent, tailored to one school, with that self-correction rule already baked in. Phase 2 is what the agent actually does with it.
Fetching agent - phase 2
The fetch agent produced by schedule-research agent only knows about its own school (a webpage to scrape). Its whole job is: given a class id, follow the recorded steps, and write the result as JSON matching one shared schema:
{
"school_id": "oaktree-primary",
"school_name": "Oaktree Primary School",
"class": "3B",
"label": "3B",
"source_url": "https://school.example/plan?class=3B",
"fetched_at": "2026-01-01T12:00:00+00:00",
"days": ["Monday", "Tuesday", "Wednesday", "Thursday", "Friday"],
"periods": [{ "number": 1, "start": "08:00", "end": "08:45" }],
"entries": [
{ "day": "Monday", "period": 1, "subject": "Math", "room": "12", "teacher": "J. Smith" }
]
}
That schema is the whole contract between phases: Phase 3 never knows or cares which school a given JSON file came from, or how it was fetched. It just reads the shared shape and lays it out.
This decoupling is what makes adding a fourth, fifth, sixth school free: each is just another fetch agent, from Phase 1, producing the same shape of file. And if the site underneath a fetch agent changes shape entirely? That's exactly what self-correcting scraper process is for in the scraper agent definition.
Once a fetch agent writes a valid JSON file, Phase 3 takes over.
Rendering the PDF - phase 3
Rendering is almost boring by comparison, which is the point: a small script opens an HTML template in a headless browser, injects the JSON files, and prints the result to PDF.
The rendering script
This wasn't hand-written either, and it wasn't designed in one pass. Claude Code was asked to prepare a few different layout templates, each rendered against the same sample data so they could be compared side by side. The one picked by the user became the actual HTML template the script uses to generate the schedule PDF. The render script itself, a small Node program using Playwright, does the boring part: open that template in a headless browser, inject the fetched JSON as data the template's own code reads, and call the browser's print-to-PDF function.
Where the whole process is described?
CLAUDE.md - a project-level instructions file sitting at the root of the repo. It's the document a Claude Code session reads first, and it's where the pipeline is actually named, with a short description of each stage and a directory map showing exactly which file does what.
Example:
# School Schedule Grabber
## What this project does
Combines class timetables from one or more school websites into a single **A4
landscape PDF**, so schedules for kids in different classes (and different
schools) can be printed or viewed side by side on one page.
Pipeline: **research → per-school fetch agent → render**.
- `schedule-research` (a Claude Code subagent) visits a school's public website,
figures out how its timetable pages work, discovers the full list of classes it
offers, and writes a dedicated fetch agent for that school.
- `fetch-<school-id>` (one generated subagent per school) knows that school's site
structure and can pull the schedule for any class at that school into a small
JSON file.
- `scripts/render-pdf.mjs` takes the fetched JSON files and the HTML template in
`templates/` and renders the combined PDF with Playwright.
It's self-correcting: if a school's site changes and its `fetch-<school-id>` agent
starts failing, that agent calls `schedule-research` back (naming its own school
id) to re-analyze the site and rewrite its own instructions, then retries — no
manual re-analysis needed.
This is designed to work across **multiple schools at once** (e.g. two kids at two
different schools) as well as multiple classes at the same school.
Data fetching process results
Takeaways
If you're building something in this shape — many similar-but-not-identical sources, each changing on its own schedule — a few things generalize past this specific project:
- Have the research agent write prompts, not code. It's more robust to a source's own inconsistencies, and "regenerating" is just rewriting a markdown file.
- Give the fetch agent an explicit, named escape hatch. The self-correcting scraper process only works if the fetch agent knows exactly who to call and what to say — put that clause in every fetch agent, verbatim.
- Pick one shared data schema and never let a fetch agent deviate from it. That's what makes adding new sources incremental instead of a rewrite.
- Treat every fetched page as untrusted data, not instructions — and while you're at it, verify your own agents' claimed side-effects instead of trusting the report.
- When something breaks, fix the shared instructions, not just the one broken run. A one-off patch fixes today's source; a template fix prevents tomorrow's version of the same bug.
Comments
Post a Comment