// case file 02 · AI automation

Lead Research Agent

Turns “find 10 B2B SaaS companies that need automation” into qualified leads with evidence, and review-ready outreach it can never send itself.

RoleSolo · design to deployTeamSoloTimelineSep 2026 · ~1 weekStackClaude Agent SDK + Sonnet 50send tools the agent can reach
// 01 the problem

Outbound was fully manual: someone defined the persona, searched for companies, read their sites and wrote cold copy from scratch. Slow, and impossible to audit.

// architecture
Lead Research Agent · Claude Agent SDKactive
every stage is a fresh Agent SDK process with only its own tools
  1. Brief + target (human)
  2. Brief complete? (decision)
  3. Stopped: names what’s missing (failure)
  4. Propose ICP: propose_icp (AI)
  5. Approve + freeze ICP (human)
  6. Discover: Apify · ≤6 calls · ≤75 (AI)
  7. Map + scrape: Firecrawl · ≤90 (AI)
  8. Redact emails/phones (step)
  9. Qualify: status · confidence · sources (AI)
  10. Target met? (decision)
  11. Recount from DB (step)
  12. Review leads (human)
  13. Draft outreach: 3 emails + LinkedIn · no send tool (AI)
  14. Admin approval (human)
  15. Export (output)
  16. Supabase log: run · tool_call · spend_ledger (store)
// 02 how it works
  1. Step 01

    Operator enters an objective and a lead target; a brief missing company type or location is stopped and told what’s missing.

  2. Step 02

    The agent proposes an ICP, labelling every value “From your brief”, “Assumed” or “Your edit”, then it’s approved and frozen.

  3. Step 03

    Discovery via Apify (LinkedIn + Google), capped at 6 calls and 75 candidates, each call with its own charge cap.

  4. Step 04

    Firecrawl maps each site first, then reads homepage, careers and about, never guessing a URL, max 90 scrapes a run.

  5. Step 05

    Each lead gets status, confidence, fit reasons, concerns and source URLs; research stops the moment the target is met.

  6. Step 06

    Drafts exist only for approved leads: a 3-step email sequence plus a LinkedIn note. An admin approves or sends back with a reason.

// numbers I’d defend
7 / 7qualified leads: Sonnet 5 vs Haiku 4.5 on the same brief
2.4×cheaper per lead on Haiku, but it rejected nobody
6 / 25Haiku claim-flags that were real (Sonnet: 3 / 3)
90+green tests while the agent never ran
// 03 decisions that mattered
01

A green test suite hid an agent that never ran

90+ tests passed while every stage route returned 503. Now a task isn’t done until the real UI has driven the real backend once.

02

Lead count is a hard cap where the money is spent

With target 3, four companies got researched. research_company now refuses with TARGET_MET and in-flight calls hold slots, so parallel calls can’t overshoot.

03

The agent’s word never decides a status

A “qualified” verdict citing a page that was never fetched is downgraded. Counts are recomputed from the database, and disagreements are stored.

04

Scraped text is data, never instructions

Emails and phones are redacted before the model sees a page, and a forged closing tag test got the envelope escaped. No tool can act on what a page says.

// incident

A green test suite hid an agent that never ran

90+ tests passed. Every stage route returned 503, and nothing reachable from a browser had ever called the agent engine.

The stage function took the runner as an injected callback and never called the real one, so the tests and the smoke script were exercising a stub.

Fix: A task isn’t done until the real UI has driven the real backend once, and anything cited as evidence gets mutation-checked (break it, watch the test fail).

// 04 results
15lead ceiling, enforced in code
0tools that can send an email
2.4×cheaper on Haiku, and still rejected
// 05 takeaway

Skills shape judgment; tools enforce rules. If the agent ignored every skill, the handlers would still refuse.

Claude Agent SDKSonnet 5Haiku 4.5ApifyFirecrawlNext.jsNode.jsSupabaseVercelRender
// what I’d do differently
  • Get one real run through end to end on day one, even with an ugly interface, then harden.
  • Test against the declared tool interface directly: eleven tools were declared and five implemented, and no test noticed.
  • Mutation-check any test I’m going to cite as evidence.
// artefacts
// contact

Let’s talk

Open to software engineering and AI automation roles, plus selective contract work. I read every message.

psst, click it
● DRAFT

Show me the copy-paste.
I’ll tell you if it should be a system.

01
02
How often?
03
Which tools?

Hiring? Tell me about the team.
I’ll tell you where I’d help first.

01
02
03
Timeline
04

You get a reply within 24 hours: whether it should be automated, what I’d use, and a rough timeline.

You get a reply within 24 hours.