Matching 3,000+ freight documents a day to the right record
Purchase orders, bills of lading and proofs of delivery arrived in every format from hundreds of counterparties. I built the multi-agent system that matched each one to the right record, and knew when to hand the hard ones to a person.
- Role
- Designed and led
- Team
- Directed a junior engineer on evaluation
- Scale
- 3,000+ documents a day, later ~5,000+ across all types
- Built
- MVP in ~3 weeks
The mess
Billing waited on a team of ~20–25 people reading every delivery document by hand. The backlog ran ~1–2 months, and whatever the documents knew stayed locked inside them.
The question
Is this document chain complete, consistent and safe to bill?
The call
The most convincing clue on a document was often the misleading one. So the system never trusted a single identifier: it had to find agreement across several before acting. We started cautious, and the billing team’s corrections decided how much more automation it earned.
What changed
- ~70% of matching automated
- False matches ~20–30% → ~5%
- Backlog ~1–2 months → ~1 week
- Review team ~20–25 → ~2–3
- It then became the document backbone other teams built on.
In hindsight
I’d build automated evaluation earlier. Manual review got us to production; it wouldn’t have kept us honest at scale.
“Document AI lives or dies on the exception path. Its real value shows up when the rest of the business can ask what it read.”
Delivered within a prior employer’s environment. Details are intentionally anonymized. I’m glad to discuss the operating problem, design principles, and production tradeoffs.
Have a workflow that looks like this?
Let’s work out whether AI is actually the right lever.