AI Operations Reporting
Pulls operations data from three sources on a schedule, computes the numbers in code, and has Claude explain them, then emails a digest you can question.
Ops numbers lived across a spreadsheet, an Airtable base and an API. Someone stitched them together by hand each period and wrote the commentary from memory.
- Schedule (trigger)
- Report config (step)
- Google Sheets (step)
- Airtable (step)
- HTTP source (step)
- Merge sources (step)
- Resolve period + normalise (step)
- Compute 26 KPIs: in code, not the model (step)
- Hash vs last run (step)
- Changed? (decision)
- Touch checked_at (store)
- Build Claude input (step)
- Claude insight: interprets only (AI)
- Parse insight (step)
- Write report (store)
- Write metrics (store)
- Write insight (store)
- Write clean records (store)
- Merge writes (step)
- Finalize run: the finish line (store)
- Email digest (output)
- Retry webhook (trigger)
- Ask webhook (trigger)
- Load report + insight + metrics (step)
- Build context (step)
- Needs model? (decision)
- Ask Claude (AI)
- Error response (failure)
- Workflow error (trigger)
- Format alert (step)
- Email operator (output)
- Step 01
A schedule trigger reads Google Sheets, Airtable and an HTTP source, then merges and normalises records.
- Step 02
The reporting period is resolved and 26 KPIs are computed in code. The model does no maths.
- Step 03
Claude gets the computed metrics and writes the insight; the response is parsed before anything is stored.
- Step 04
Report, metric rows and insight are written to Supabase, then a digest is emailed.
- Step 05
A hash is compared with the last completed run; if nothing changed, it just stamps checked_at instead of re-running.
- Step 06
A webhook lets someone ask a question about a report, answered from the stored report and metrics.
Numbers from code, words from the model
Claude gets a compact digest of the computed numbers plus data-quality flags, and interprets them. It never calculates.
Change detection before spend
If the sources haven’t moved since the last completed run, the Claude call and the writes are skipped entirely.
Every write is a stored artefact
Reports, metrics and insights land as rows, so any digest can be traced back.
Three writes, no finish line
The report, its metrics and its insight were written as three independent steps.
If one died halfway, a half-written report became the newest one, and nothing said so.
Fix: "This run is done" is its own deliberate finalisation step, and a report without it is never shown as current.
Let the model do language, not arithmetic.
- List every failure class first (a bad row, a dead source, a source returning rubbish, the model down, a failed write, a failed email, a schedule that never fires) and give each its own behaviour and test before the happy path.
- Keep every number and quality flag in code, and lock the prompt to interpret-only.