Lead Research Agent
Turns “find 10 B2B SaaS companies that need automation” into qualified leads with evidence, and review-ready outreach it can never send itself.
Outbound was fully manual: someone defined the persona, searched for companies, read their sites and wrote cold copy from scratch. Slow, and impossible to audit.
- Brief + target (human)
- Brief complete? (decision)
- Stopped: names what’s missing (failure)
- Propose ICP: propose_icp (AI)
- Approve + freeze ICP (human)
- Discover: Apify · ≤6 calls · ≤75 (AI)
- Map + scrape: Firecrawl · ≤90 (AI)
- Redact emails/phones (step)
- Qualify: status · confidence · sources (AI)
- Target met? (decision)
- Recount from DB (step)
- Review leads (human)
- Draft outreach: 3 emails + LinkedIn · no send tool (AI)
- Admin approval (human)
- Export (output)
- Supabase log: run · tool_call · spend_ledger (store)
- Step 01
Operator enters an objective and a lead target; a brief missing company type or location is stopped and told what’s missing.
- Step 02
The agent proposes an ICP, labelling every value “From your brief”, “Assumed” or “Your edit”, then it’s approved and frozen.
- Step 03
Discovery via Apify (LinkedIn + Google), capped at 6 calls and 75 candidates, each call with its own charge cap.
- Step 04
Firecrawl maps each site first, then reads homepage, careers and about, never guessing a URL, max 90 scrapes a run.
- Step 05
Each lead gets status, confidence, fit reasons, concerns and source URLs; research stops the moment the target is met.
- Step 06
Drafts exist only for approved leads: a 3-step email sequence plus a LinkedIn note. An admin approves or sends back with a reason.
A green test suite hid an agent that never ran
90+ tests passed while every stage route returned 503. Now a task isn’t done until the real UI has driven the real backend once.
Lead count is a hard cap where the money is spent
With target 3, four companies got researched. research_company now refuses with TARGET_MET and in-flight calls hold slots, so parallel calls can’t overshoot.
The agent’s word never decides a status
A “qualified” verdict citing a page that was never fetched is downgraded. Counts are recomputed from the database, and disagreements are stored.
Scraped text is data, never instructions
Emails and phones are redacted before the model sees a page, and a forged closing tag test got the envelope escaped. No tool can act on what a page says.
A green test suite hid an agent that never ran
90+ tests passed. Every stage route returned 503, and nothing reachable from a browser had ever called the agent engine.
The stage function took the runner as an injected callback and never called the real one, so the tests and the smoke script were exercising a stub.
Fix: A task isn’t done until the real UI has driven the real backend once, and anything cited as evidence gets mutation-checked (break it, watch the test fail).
Skills shape judgment; tools enforce rules. If the agent ignored every skill, the handlers would still refuse.
- Get one real run through end to end on day one, even with an ugly interface, then harden.
- Test against the declared tool interface directly: eleven tools were declared and five implemented, and no test noticed.
- Mutation-check any test I’m going to cite as evidence.