Invoices and accounts payable
Line item extraction, purchase order matching and exception handling, connected to the finance system you already run.
Most document automation is sold on adjectives. We sell it on numbers. Oligamy Software builds IDP pipelines that classify, extract and validate data from invoices, contracts, claims and scanned archives, then reports field level accuracy, cost per thousand pages and latency against your current baseline.
IDP automates reading business documents end to end: it classifies each file, extracts the fields that matter, validates them against business rules, and passes structured data into the systems that need it. It combines OCR, layout models and language models rather than relying on any single technique.
PDFs, scans, emails and images from every source you receive them on.
Decide what each file is, and split multi document PDFs before anything else runs.
Pull the named fields that matter for this document class, with a confidence score attached.
Check values against business rules. Anything below the threshold goes to a human, not to production.
Structured data lands in the ERP, CRM or claims platform you already run.
OCR converts an image of text into characters and stops there. OCR is one component of IDP, not a substitute for it.
| Capability | OCR alone | Full IDP |
|---|---|---|
| Turns an image of text into characters | Yes | Yes |
| Recognises what kind of document it is | No | Yes |
| Extracts specific named fields | No | Yes |
| Understands layout and context | No | Yes |
| Validates values against business rules | No | Yes |
| Flags low confidence cases for a human | No | Yes |
| Feeds structured data into your systems | No | Yes |
There is no single best approach to document extraction, only a best approach per document class. Vendors publish definitions; almost nobody publishes where their method fails. This is the comparison we run before proposing anything.
| Approach | Accuracy | Cost per page | Where it breaks |
|---|---|---|---|
| Classic OCR plus rules | High on clean, fixed layouts | Lowest | Any layout it has not seen before |
| Cloud document AI | Strong on common forms | Low to medium | Domain specific fields, rare languages |
| Layout model | Strong on semi structured | Medium | Needs labelled data to get there |
| LLM, zero shot | Flexible but inconsistent | Highest at volume | Silent invention on missing fields |
| LLM with constrained schema | High and predictable | Medium to high | Needs a schema per document class |
| Hybrid pipeline | Highest measured | Tuned per class | More moving parts to maintain |
The bars are our engineering judgement on a five point scale, not the result of a benchmark we ran for you. Absolute numbers only mean something against a specific document set, so the measured figures come from the audit on your own files. A longer cost bar means more expensive.
A pipeline that cannot tell you when it is unsure is not automation. It is a faster way to be wrong.
Every pipeline we ship has a confidence threshold and a human path behind it. That is the difference between a system you can put in front of an auditor and a demo.
Find your document class on the left, read across. The pattern is the argument: classic OCR collapses the moment documents stop being clean and fixed, and the approaches that hold up on hard classes are the ones that cost the most to run. That is the trade you are actually making.
| Document class | OCR + rules | Cloud doc AI | Layout model | LLM zero shot | LLM + schema | Hybrid |
|---|---|---|---|---|---|---|
| Clean digital invoices | OCR + rules: Strong, 4 out of 5 | Cloud doc AI: Best fit, 5 out of 5 | Layout model: Best fit, 5 out of 5 | LLM zero shot: Workable, 3 out of 5 | LLM + schema: Best fit, 5 out of 5 | Hybrid: Best fit, 5 out of 5 |
| Semi structured forms | OCR + rules: Limited, 2 out of 5 | Cloud doc AI: Strong, 4 out of 5 | Layout model: Best fit, 5 out of 5 | LLM zero shot: Workable, 3 out of 5 | LLM + schema: Strong, 4 out of 5 | Hybrid: Best fit, 5 out of 5 |
| Contracts, long and varied | OCR + rules: Poor fit, 1 out of 5 | Cloud doc AI: Limited, 2 out of 5 | Layout model: Workable, 3 out of 5 | LLM zero shot: Strong, 4 out of 5 | LLM + schema: Best fit, 5 out of 5 | Hybrid: Best fit, 5 out of 5 |
| Insurance claim bundles | OCR + rules: Poor fit, 1 out of 5 | Cloud doc AI: Limited, 2 out of 5 | Layout model: Workable, 3 out of 5 | LLM zero shot: Workable, 3 out of 5 | LLM + schema: Strong, 4 out of 5 | Hybrid: Best fit, 5 out of 5 |
| Handwriting and poor scans | OCR + rules: Poor fit, 1 out of 5 | Cloud doc AI: Workable, 3 out of 5 | Layout model: Limited, 2 out of 5 | LLM zero shot: Workable, 3 out of 5 | LLM + schema: Workable, 3 out of 5 | Hybrid: Strong, 4 out of 5 |
| Non English documents | OCR + rules: Limited, 2 out of 5 | Cloud doc AI: Workable, 3 out of 5 | Layout model: Workable, 3 out of 5 | LLM zero shot: Strong, 4 out of 5 | LLM + schema: Strong, 4 out of 5 | Hybrid: Best fit, 5 out of 5 |
Ratings are our engineering judgement on a five point scale, not a benchmark we ran for you. They are the starting hypothesis; the audit on your own files replaces them with measured numbers.
Six document classes we see most often, and what changes when each one goes through a pipeline instead of a person.
Line item extraction, purchase order matching and exception handling, connected to the finance system you already run.
Claim intake, supporting document validation and structured handoff to the claims platform, with a full audit trail.
Clause and obligation extraction across templates that were never standardised, including scanned amendments.
Handwriting, low quality faxes, multi column layouts and non English documents, where off the shelf OCR degrades first.
Routing mixed batches and multi document PDFs to the right workflow before anything is extracted.
Turning a document archive into a system that answers questions with citations rather than plausible guesses.
Four stages, and the first one is measurement. We do not quote before we know what your current process gets right.
We take a representative sample of your real documents and measure what your current extraction actually gets right. That number becomes the target the new pipeline has to beat.
We pick the approach per document class from the table above, define the validation rules, and set the confidence threshold that sends a document to a human instead of guessing.
We benchmark against your baseline on field level precision and recall, cost per thousand pages and p95 latency, before anything touches production.
We integrate with the ERP, CRM or claims platform you already run, ship with a rollback path, and keep watching accuracy drift as document formats change.
Taken from what people search alongside this topic, not from what we wish they asked.
Intelligent document processing (IDP) automates reading business documents end to end: it classifies each file, extracts the fields that matter, validates them against business rules, and passes structured data into the systems that need it. It combines OCR, layout models and language models rather than relying on any single technique.
OCR converts an image of text into characters and stops there. IDP is the full pipeline around it: deciding what kind of document it is, understanding layout and context, pulling out specific fields, checking them for plausibility, and routing anything uncertain to a human. OCR is one component of IDP, not a substitute for it.
Accuracy depends on document class, not on the vendor's marketing. Clean structured invoices routinely reach very high field level accuracy, while handwriting, poor scans and unusual layouts drop sharply. Oligamy Software benchmarks accuracy on a sample of your real documents before proposing an approach, and reports precision and recall per field rather than a single headline number.
Cost per page varies by an order of magnitude depending on the approach: classic OCR with rules is the cheapest per page but the most brittle, general purpose vision models are the most flexible but the most expensive at volume, and hybrid pipelines sit in between. Oligamy Software measures cost per thousand pages alongside accuracy so the trade off is explicit rather than assumed.
Yes, but these are exactly the cases where off the shelf tools degrade. Handwriting, low quality faxes, multi column layouts and non English documents need a pipeline designed for them, usually with a lower confidence threshold and an explicit human review path. Oligamy Software treats these as their own document classes with their own measured accuracy.
Those are the workflows where it pays off most, because document volume is high and a wrong field is a compliance incident rather than a typo. It requires validation rules, confidence thresholds, a human review path for uncertain extractions, and a full audit trail of what the system decided and why. Oligamy Software builds document pipelines for regulated fintech and insurance workflows with those controls in place.
Send a sample of your real files. We measure what your current process gets right, benchmark the alternatives against it, and give you the numbers whether or not you build with us.
Get a document audit