// case file 05 · AI automation

Intelligent Invoice Processing

Watches an inbox, pulls invoice fields out of any PDF (even a damaged scan) and logs them once, with failures recorded instead of dropped.

RoleSolo · workflow designTeamSoloTimelineAug 2026Stackn8n + Gmail2extraction paths: text + OCR
// 01 the problem

Invoices arrived as email attachments in every format (clean PDFs, scans without a text layer, broken files) and were keyed in by hand.

// architecture
Intelligent Invoice Processing · n8nactive
  1. Watch inbox (trigger)
  2. Prepare & triage (step)
  3. Invoice? (decision)
  4. Has PDF? (decision)
  5. Split attachments (step)
  6. Merge candidates (step)
  7. Per-email loop (step)
  8. Has binary? (decision)
  9. Extract PDF text (step)
  10. Has text? (decision)
  11. Adopt text (step)
  12. Mistral OCR (AI)
  13. OCR text? (decision)
  14. PDF extraction error (failure)
  15. Empty extraction failure (failure)
  16. Finalize failure row (failure)
  17. Consolidate invoice text (step)
  18. Pre-dedupe anchor (step)
  19. Already failed? (decision)
  20. Restore failure fields (step)
  21. Claude extracts fields (AI)
  22. Validate & assemble (step)
  23. Has invoice #? (decision)
  24. Check dedupe (step)
  25. Duplicate? (decision)
  26. Log invoice (store)
  27. Mark as read (output)
// 02 how it works
  1. Step 01

    A Gmail trigger watches the inbox and filters for likely invoices with PDF attachments.

  2. Step 02

    Each attachment is split out; text is extracted directly, or routed to Mistral OCR when there’s no text layer.

  3. Step 03

    Claude extracts structured invoice fields from the consolidated text.

  4. Step 04

    The invoice number is checked against the sheet so the same invoice is never logged twice.

  5. Step 05

    Valid rows are logged to Google Sheets; extraction failures become their own failure rows.

  6. Step 06

    The email is marked read once it’s been handled.

// numbers I’d defend
31nodes
6test invoices incl. a damaged scan
2extraction paths
// 03 decisions that mattered
01

Failures are rows, not silence

An empty extraction or broken PDF is logged with a reason, so nothing disappears.

02

OCR only when needed

Direct text extraction first; OCR is the fallback for scans.

03

Dedupe against the source of truth

The sheet itself is the authority on what’s already been logged.

// 04 results
31nodes in the workflow
2extraction paths: text + OCR
6test invoices incl. a damaged scan
// 05 takeaway

Test with the ugliest real inputs first: the damaged scan is the spec.

n8nGmailClaudeMistral OCRGoogle Sheets
// what I’d do differently
  • Confirm a feature is needed against the real test cases before building it. A fallback dedupe for invoices without numbers took several bug fixes, then added nothing, because those rows were already flagged invalid.
// artefacts
// contact

Let’s talk

Open to software engineering and AI automation roles, plus selective contract work. I read every message.

psst, click it
● DRAFT

Show me the copy-paste.
I’ll tell you if it should be a system.

01
02
How often?
03
Which tools?

Hiring? Tell me about the team.
I’ll tell you where I’d help first.

01
02
03
Timeline
04

You get a reply within 24 hours: whether it should be automated, what I’d use, and a rough timeline.

You get a reply within 24 hours.