Intelligent Invoice Processing
Watches an inbox, pulls invoice fields out of any PDF (even a damaged scan) and logs them once, with failures recorded instead of dropped.
Invoices arrived as email attachments in every format (clean PDFs, scans without a text layer, broken files) and were keyed in by hand.
- Watch inbox (trigger)
- Prepare & triage (step)
- Invoice? (decision)
- Has PDF? (decision)
- Split attachments (step)
- Merge candidates (step)
- Per-email loop (step)
- Has binary? (decision)
- Extract PDF text (step)
- Has text? (decision)
- Adopt text (step)
- Mistral OCR (AI)
- OCR text? (decision)
- PDF extraction error (failure)
- Empty extraction failure (failure)
- Finalize failure row (failure)
- Consolidate invoice text (step)
- Pre-dedupe anchor (step)
- Already failed? (decision)
- Restore failure fields (step)
- Claude extracts fields (AI)
- Validate & assemble (step)
- Has invoice #? (decision)
- Check dedupe (step)
- Duplicate? (decision)
- Log invoice (store)
- Mark as read (output)
- Step 01
A Gmail trigger watches the inbox and filters for likely invoices with PDF attachments.
- Step 02
Each attachment is split out; text is extracted directly, or routed to Mistral OCR when there’s no text layer.
- Step 03
Claude extracts structured invoice fields from the consolidated text.
- Step 04
The invoice number is checked against the sheet so the same invoice is never logged twice.
- Step 05
Valid rows are logged to Google Sheets; extraction failures become their own failure rows.
- Step 06
The email is marked read once it’s been handled.
Failures are rows, not silence
An empty extraction or broken PDF is logged with a reason, so nothing disappears.
OCR only when needed
Direct text extraction first; OCR is the fallback for scans.
Dedupe against the source of truth
The sheet itself is the authority on what’s already been logged.
Test with the ugliest real inputs first: the damaged scan is the spec.
- Confirm a feature is needed against the real test cases before building it. A fallback dedupe for invoices without numbers took several bug fixes, then added nothing, because those rows were already flagged invalid.