Document Extraction Accuracy: A Complete Field-Level Evaluation Guide for Invoice AI & OCR
Document extraction is the process of pulling structured fields — an invoice number, a total amount, a due date — out of unstructured documents, and a single "overall accuracy" number for that process is almost always misleading, because it can be 97% correct on every field that doesn't matter and still be catastrophically wrong on the one field a payment run depends on. 📄
The stakes aren't abstract. A finance team that ships an extraction model at "96% accuracy" and later discovers that the 4% of errors are concentrated in the "Amount Due" field — not scattered evenly across vendor name and page number — has, in practice, built a system that quietly overpays or underpays suppliers on a predictable schedule. Field-level, business-aware evaluation is what tells you which 4% you're looking at before it costs money. 💸
One aggregate accuracy score can't watch five different doors at once — each stage of the pipeline needs its own check.
📑 In This Post
- Why One Accuracy Score Can't Grade a Document
- The Stages of a Document AI Pipeline
- Building a Representative Test Set
- The Metric Suite: Field, Exact-Match, Rules, Confidence, Document Pass Rate
- The Business Risks Hiding Inside a Good-Looking Score
- Why Valid JSON Isn't the Same as a Correct Invoice
- When to Automate, When to Ask a Human
- What's Changed Recently: Straight-Through Processing as the KPI
- Rolling This Out at Enterprise Scale
- A Worked Invoice Example With a Scorecard
- Common Mistakes
- Practical Deployment Checklist
- FAQ
- References & Further Reading
🔀 Quick Comparison: The Metrics That Actually Belong in an Extraction Eval
| Metric | What it tells you | What it misses |
|---|---|---|
| Document-level accuracy | Rough sense of overall model health across a batch. | Which specific field is failing, and whether that field is the one that moves money. |
| Field-level accuracy | Per-field pass rate — e.g. supplier name 98%, amount due 91%. | Whether an error is a rounding difference or a completely wrong value. |
| Exact-value match | Strict string/number equality against ground truth, after normalization. | Whether the number is plausible in context (a validation rule's job). |
| Validation-rule pass rate | Whether extracted values are internally consistent (line items sum to total, dates are sane). | Whether an individually-plausible value is still the wrong one. |
| Confidence calibration | Whether the model's own certainty score can be trusted to route work. | Correctness itself — a confident model can be confidently wrong. |
| Document-level pass rate | The metric closest to business reality: "did every critical field on this document clear?" | Diagnosis — it tells you something is wrong but not which field or why. |
1. Why One Accuracy Score Can't Grade a Document
Kid Analogy
imagine a spelling test with ten words, and a teacher who only tells you "you got 90% right" without saying which word you missed. If the missed word was "cat," no big deal. If the missed word was your own name on the top of the paper, that single miss matters far more than the other nine correct words combined. Grading a document extraction system the same way — one blended percentage — makes the same mistake at a much larger scale. 🐱
A typical invoice has ten to thirty extractable fields: header fields (invoice number, purchase order number, supplier name, dates, totals, tax) and repeating line-item fields (description, quantity, unit price, line amount). An "overall accuracy" figure — often computed as correct-fields-divided-by-total-fields across a whole test set — treats a mistake on "supplier fax number" exactly the same as a mistake on "amount due." Both fields count once. Both contribute equally to the average. But only one of them determines how much money leaves the bank account.
This is a well-documented distinction in production intelligent document processing (IDP) tooling, not a theoretical concern. Microsoft's own Azure AI Document Intelligence documentation draws an explicit line between two different numbers reported for a custom extraction model: an estimated accuracy per field, calculated by testing the trained model against held-out combinations of its own training data, and a per-prediction confidence score returned at inference time for each individual field on each individual document.[1] Those are not the same measurement, and conflating them is one of the most common evaluation mistakes teams make — a model can have high estimated accuracy on paper while still returning a low-confidence, likely-wrong "AmountDue" on the one invoice in front of you right now.
✅ Worked example: a shared-services accounts-payable team measuring "94% field accuracy" across 500 test invoices discovers, once they break the number down by field, that supplier name and invoice date are both above 99% — but "Amount Due" sits at 88% and "Tax Amount" sits at 81%. The blended 94% never would have surfaced that the two fields controlling the actual payment are the weakest ones in the system.
💡 Harder case: now suppose "Amount Due" is 97% correct — but on the 3% of invoices where it's wrong, the model isn't slightly off, it's reading the subtotal before tax instead of the final total. A field-level percentage alone won't tell you that the errors are systematic and concentrated on multi-currency invoices from one template family, which is exactly the pattern that turns into a real financial loss instead of noise.
🎯 Use this when: you're asked to sign off on a document extraction system and the only number you've been handed is a single accuracy percentage. Ask for the field-level breakdown before you approve anything.
2. The Stages of a Document AI Pipeline
Kid Analogy
think of a document like a letter passing through a relay race — one runner scans it, one reads the handwriting, one figures out which word means what, one double-checks the numbers add up, and the last one calls a grown-up if something looks off. If any single runner drops the baton, everyone after them is working with bad information, even if they run their own leg perfectly. 🏃
A production document AI pipeline is almost always five distinct stages, and each one needs its own evaluation, because errors introduced early are invisible — and unfixable — downstream.
- Document input — the file arrives (email attachment, scanner, upload portal, EDI feed). This stage decides file type handling, page ordering, and whether a multi-invoice PDF gets correctly split into one document per invoice. A pipeline that silently merges two invoices into one "document" corrupts everything after it.
- OCR or vision processing — pixels become text (classic OCR) or the page image is handed directly to a vision-capable model. This is where skew, low contrast, and handwriting cause character-level errors — a "0" read as an "8," a decimal point dropped.
- Extraction — text and layout are mapped onto a schema: which span of text is the "invoice number" versus the "purchase order number." This is where a technically-correct piece of text gets attached to the wrong label.
- Validation — extracted values are checked against business rules independent of the model: do line items sum to the subtotal, is the due date after the invoice date, does the supplier exist in the vendor master.
- Human review — a person resolves anything that failed validation, fell below a confidence threshold, or hit a business rule (e.g., first invoice from a new vendor, or amount above an approval limit).
Amazon's own Textract documentation reflects this same staged thinking in its production guidance: it recommends attaching a confidence score to every extracted element and routing anything below a use-case-specific threshold to a human reviewer rather than trusting a blended pass/fail at the end.[2] AWS's published best practices are explicit that the right threshold is not fixed — the documentation gives archival or handwritten-note use cases as low as roughly 50% confidence tolerance, identity-document fields around 80%, and financial decision processes at 90% or higher, because the cost of a wrong answer is different in each case.[2]
✅ Worked example, continued: in our accounts-payable team's pipeline, "Amount Due" and "Tax Amount" being the weakest fields (from Section 1) turns out to trace back to Stage 2 — OCR — not Stage 3. Their poorly-scanned supplier invoices have currency symbols and decimal points that get misread before extraction ever sees the text, so no amount of extraction-model tuning would have fixed it. Staged evaluation is what located the actual broken runner in the relay.
🎯 Use this when: an extraction error shows up and the instinct is to "retrain the extraction model." Check which stage actually produced the error first.
3. Building a Representative Test Set
Kid Analogy
if a driving test only ever happened on a sunny day, on an empty, perfectly straight road, passing it wouldn't tell you whether someone can actually drive — it would only tell you they can drive in the one condition you tested. A document test set has the same problem: if it's all clean, upright, single-template PDFs, a 98% pass rate is measuring the easy 80% of real traffic and saying nothing about the hard 20%. 🚗
A representative test set for document extraction deliberately includes the conditions that are rare in a tidy internal demo but common in a live inbox:
- Clean, digitally-generated documents — the baseline; if the system can't handle these near-perfectly, nothing else matters.
- Poor scans — low resolution, faxed, photocopied-of-a-photocopy, uneven lighting from a phone camera.
- Rotated or skewed images — pages scanned upside down, sideways, or at an angle.
- Tables and line items — multi-row, multi-column layouts, including merged cells and tables that span page breaks.
- Handwriting — handwritten totals, signatures, or annotations layered on a printed form.
- Multiple templates — the same field (say, "invoice number") appearing in a different position, label wording, and font across dozens of supplier layouts.
- Missing or blank fields — documents where a field genuinely isn't present, so the correct extraction is "null," not a guess.
That last category is easy to skip and important not to: a model graded only on documents where every field exists learns nothing about whether it can correctly say "I don't know" instead of hallucinating a plausible-looking value into an empty cell.
💡 Harder case: our accounts-payable team's original 500-document test set was pulled entirely from documents already successfully processed in the old system — meaning it systematically excluded the poor scans and unusual templates that historically caused manual intervention. The test set was, without anyone intending it, pre-filtered to the easy cases. Rebuilding it to include a proportional share of past manual-review documents dropped the measured field accuracy on "Amount Due" from a comfortable 96% to a more honest 88% — the same underlying model, a materially different truth.
🎯 Use this when: a vendor or internal team reports a headline accuracy number. Ask what fraction of the test set was poor scans, rotated pages, and missing-field documents before trusting it.
4. The Metric Suite: Field Accuracy, Exact-Match, Validation Rules, Confidence, and Document Pass Rate
Kid Analogy
a doctor's checkup doesn't rely on one number either — it's temperature, blood pressure, pulse, and a few others together, because any single vital sign can look normal while something else is wrong. A document extraction system needs the same kind of multi-vital-sign checkup instead of one thermometer reading. 🩺
Five metrics, used together, cover the ground a single score can't:
- Field-level accuracy — per-field correctness across the test set (invoice number: 99%, amount due: 91%, line-item description: 85%), so weak fields are visible instead of averaged away.
- Exact-value match — strict equality between the extracted value and ground truth after a defined normalization step (currency symbols stripped, dates in one format, whitespace trimmed) — this is the honest "did we get the actual value right," as opposed to a fuzzy string-similarity score that can call "$1,240.00" and "$1,204.00" mostly-the-same when they are two very different amounts.
- Validation-rule pass rate — independent business-logic checks run on the extracted output itself: do the line items sum to the subtotal, does subtotal plus tax equal the total, is the due date on or after the invoice date, is the invoice number in the expected format for that supplier. These checks catch errors that field-level accuracy can miss entirely, because a value can be individually plausible and still be wrong in context.
- Confidence thresholds — the model's own self-reported certainty, evaluated for calibration: if a field is scored at 95% confidence, does it actually turn out to be right about 95% of the time on your own documents? Microsoft's Document Intelligence guidance builds its confidence-score explanation around this same "does the stated probability hold up" framing, and recommends checking it empirically before treating a threshold as an automation gate.[1]
- Document-level pass rate — the percentage of whole documents where every critical field cleared its checks, which is the number closest to "how much of this batch can go through untouched." This is intentionally the strictest metric: one failed critical field fails the whole document, because that's how accounts payable actually works — a wrong amount on an otherwise-perfect invoice still blocks payment.
✅ Worked example, continued: once the accounts-payable team added a validation rule — "sum of line-item amounts must equal subtotal, within one cent" — it caught a class of error field-level accuracy had missed entirely: the model was correctly reading each individual line-item amount, but occasionally assigning one line's dollar figure to the row above or below it. Every individual number was "correct" somewhere on the page; the validation rule was what proved it was attached to the wrong row.
🎯 Use this when: designing an eval harness for any extraction system — build all five metrics before writing a single line of model-tuning code, not after.
5. The Business Risks Hiding Inside a Good-Looking Score
Kid Analogy
mixing up two friends' names by accident is embarrassing; mixing up two friends' lunch money is a problem. Some extraction mistakes are the embarrassing kind and some are the lunch-money kind, and an evaluation plan has to know which is which before it decides what "good enough" means. 💰
Three failure patterns recur across document extraction deployments, and all three can hide behind a healthy-looking aggregate score:
- Wrong payment amount — the model reads the subtotal, the tax, or a line-item total instead of the grand total, or misreads a digit ("1,340.00" as "1,840.00"). This is the single highest-consequence error class because it directly changes how much money moves.
- Invoice number confused with purchase order number — both are often similarly-formatted alphanumeric strings on the same page, sometimes in adjacent fields. Swapping them breaks three-way matching (PO, receipt, invoice) downstream and can cause a payment to be matched to the wrong purchase order entirely, or rejected and stalled in a queue.
- Value assigned to the wrong line item — as in the worked example above, a correct number attached to the wrong row. This is especially dangerous because every individual value looks right in isolation; only row-level, positional checking catches it.
💡 Harder case: a purchase-order/invoice-number swap can pass every metric described in Section 4 and still cause real damage. The swapped value is a real string that exists somewhere on the document (so exact-match against a mislabeled ground-truth field could even look "correct" if the test-set labeling itself has the same confusion baked in), it's the right data type and format (so a validation rule checking "is this alphanumeric" won't catch it), and confidence can be high (the OCR read the text perfectly — the field is just mislabeled). Catching this specific failure mode usually requires a dedicated validation rule that cross-checks the extracted PO number against a live purchase-order system, not just internal consistency.
🎯 Use this when: prioritizing which fields get the strictest thresholds and the most test-set coverage — rank fields by financial and matching consequence, not by how hard they are to extract.
6. Why Valid JSON Isn't the Same as a Correct Invoice
Kid Analogy
a book report that's neatly typed, correctly spelled, and turned in on time can still get the plot completely wrong. Looking right and being right are two different checks, and grading only the first one misses the point of the assignment. 📝
Modern extraction systems, especially those built on large language models, are generally reliable at producing well-formed output — a JSON object with the right keys, correctly typed values, no missing braces. That reliability is genuinely useful, and it's also completely orthogonal to whether the values inside that JSON are correct. A model can return
{
"invoice_number": "INV-88213",
"amount_due": 1840.00,
"due_date": "2026-11-02"
}
— perfectly valid, perfectly typed JSON — while the real amount due on the source document is $1,340.00. Schema validation (does this parse, are the types right) is a necessary but separate check from value validation (is this the right number). Treating "the model returned valid JSON" as evidence of correctness is one of the more common blind spots in LLM-based extraction pipelines specifically, because structured-output features make schema compliance close to automatic, which makes it tempting to assume the hard part is done.
This is exactly why the validation stage from Section 2 has to run independently of, and after, extraction — it's checking a different property of the output than the extraction step itself was optimized to get right.
🎯 Use this when: reviewing an LLM-based extraction pipeline's test results — separate "percentage of outputs that parsed as valid schema" from "percentage of field values that were correct" as two distinct, both-mandatory numbers.
7. When to Automate, When to Ask a Human — and Watching for New Failure Patterns
Kid Analogy
a kid doing a jigsaw puzzle keeps going by themselves as long as pieces obviously fit, but stops and asks for help on a piece they're not sure about, rather than jamming it in and hoping. A good extraction pipeline needs the same instinct — built in, not left to chance. 🧩
The routing decision — straight-through automation versus human review — should be driven by the combination of confidence score and validation-rule outcome, not either alone:
- Automate when confidence is above the calibrated threshold for that field and every relevant validation rule passes (totals reconcile, dates are sane, format matches the supplier's known pattern).
- Route to human review when confidence falls below threshold on any critical field, a validation rule fails, the document is from a first-seen supplier or template, or the amount exceeds a defined approval limit regardless of confidence.
AWS's Textract best-practice guidance describes exactly this pattern using Amazon Augmented AI (A2I): a reviewer sets the confidence rules that decide when a low-confidence field gets kicked to a human loop, and — importantly — recommend also randomly sampling a percentage of high-confidence, auto-approved documents for review, purely to monitor whether the model is drifting.[3] That random sampling is the part that's easy to skip and expensive to skip: without it, a model that's quietly gotten worse on a specific template or vendor never surfaces, because everything above threshold sails through untouched and nobody is watching it anymore.
In production, failure-pattern monitoring means tracking, on a rolling basis: which fields are failing validation most often, which supplier templates have the lowest confidence scores, whether the human-review rate is creeping up for a specific document type, and whether corrections made by reviewers cluster around a particular error (the PO/invoice-number swap from Section 5 is a classic recurring pattern worth its own dashboard tile).
🎯 Use this when: setting up production monitoring for a live extraction system — a routing rule without a random-sample audit of the auto-approved side is a rule you can't verify is still working.
8. What's Changed Recently: Straight-Through Processing Is the KPI Now, and the "Demo Number" Isn't the Production Number
Kid Analogy
Imagine a school bragging that "90% of homework gets checked by the teacher within a day" — but that number only counts homework from the kids who always turn tidy, complete pages. The kids whose pages are smudged or missing a name still wait a week. The bragging number and the number for the hardest homework are two very different stories, and a smart parent asks for both.
Finance and AP-automation teams have largely stopped talking about "extraction accuracy" as the headline number and shifted to straight-through processing (STP) rate — the share of documents that flow from intake to posting with zero human touch. It's a natural evolution of the same document-level pass rate described in Section 4, just renamed and adopted as a company-wide KPI rather than an engineering metric, which is exactly why finance leaders and evaluation engineers now need to be speaking about the same number.
Two things are worth watching closely if this is the metric your organization is now being measured on:
- The demo-versus-production gap is real and it's large. A platform's showcase number, generated on a handful of pristine, hand-picked invoices against a freshly cleaned-up vendor list, routinely looks far better than what the same platform delivers once it's handling messy, real supplier traffic. Independent AP-industry survey work has reported typical automation adopters clustering well below the touchless-rate figures used in vendor marketing, with only a minority of programs reaching those higher numbers. The more useful question for any vendor isn't whether the platform supports touchless processing at all — it's what that number actually looks like once it's running on live volume comparable to yours, and whether they can back the claim with evidence rather than a demo. Treat any single headline percentage, ours or a vendor's, with the same skepticism.
- Exception handling, not extraction, is now the bottleneck teams talk about most. Once extraction accuracy on clean documents is reasonably solid — which it now often is — what keeps STP rate from climbing further is the same handful of things this whole post has been about: non-PO invoices, mismatched line items, missing fields, and first-seen supplier templates. The industry language for the newer generation of tools that try to close that gap is "agentic" document processing — systems that don't just extract a field but reason about an exception (why doesn't this line item match the PO?) before deciding whether it's safe to proceed. That's a genuinely useful direction, but it doesn't remove the need for the evaluation discipline in this post; if anything, an agent that can also take action on an exception needs its decisions evaluated, not just its extracted values, which is a strictly harder evaluation problem than the one this post has focused on.
For an evaluation team, the practical takeaway is to keep computing document-level pass rate exactly as described in Section 4 — but start reporting it in the same STP-rate language finance stakeholders are now using, broken out by document type and supplier segment rather than as one company-wide number, since that's where the real gap between "our demo" and "our production" tends to hide.
🎯 Use this when: a leadership team asks for "our touchless rate" — deliver it segmented by document type and supplier, not as a single blended figure, for exactly the reason Section 1 gives for field-level accuracy.
9. Rolling This Out at Enterprise Scale
Kid Analogy
one kid keeping their own homework folder organized is easy. A whole school making sure every classroom's homework folder is organized the same way, updated when the curriculum changes, and checked by someone before report cards go out — that needs actual rules, not good intentions. 🏫
Scaling document-extraction evaluation past a single team's pilot requires a handful of governance pieces that are easy to skip early and expensive to retrofit later:
- Ownership and governance — a named owner for the eval pipeline itself (not just the extraction model), responsible for the test set, the validation rules, and the routing thresholds as a maintained system.
- Test-set versioning — the test set has a version number, and every reported accuracy figure states which version it was measured against; test sets go stale as new supplier templates, currencies, or document types show up in production, so a scheduled refresh cadence matters as much as the initial build.
- CI-gated evaluation — any prompt, model, or extraction-logic change runs the full metric suite from Section 4 against the versioned test set before it ships, with defined minimum thresholds per critical field that block a release if not met.
- Access control and data governance — test sets built from real invoices contain real supplier names, amounts, and banking references; they need the same access controls as production financial data, not looser ones just because they're "test" data.
- Cost governance for LLM-as-judge or model-assisted evaluation — if part of the eval uses another model to grade extraction outputs, that's a recurring cost that scales with test-set size and evaluation frequency; budget and rate-limit it the way any other production API dependency would be.
- Dashboards separating stages — a dashboard tile for OCR-stage confidence is a different signal than one for extraction-stage field accuracy or validation-rule pass rate; collapsing them into one dashboard number recreates the exact problem this whole post is about, one layer up.
- Alerting for regressions — an automated alert when document-level pass rate or any critical field's accuracy drops below its baseline on the rolling production sample, not just at model-release time.
🎯 Use this when: an extraction pipeline moves from one team's pilot to a shared service used by multiple business units — that transition is the moment to put these governance pieces in place, before scale makes retrofitting them painful.
10. A Worked Invoice Example With a Scorecard
This section is a small, hands-on walkthrough you can copy for a real pilot — scoped to a throwaway batch of sample invoices, not production data, so it's safe to run today.
Pull 20 real (or realistic sample) invoices you're not currently using in any model training. Aim for a deliberate spread: a few clean digital PDFs, a few scanned photocopies, one or two rotated or skewed, at least three with multi-row line-item tables, and two with a field genuinely missing (no PO number, for example).
Hand-label the ground truth yourself for seven fields on each invoice: invoice number, purchase order number, supplier name, invoice date, due date, tax amount, and amount due, plus every line item's description, quantity, and amount. This is tedious and that's fine — it only has to happen once per document, and it's the foundation everything else is measured against.
Run the batch through your extraction system and build a simple spreadsheet: one row per document, one column per field, and for each cell record extracted value, ground-truth value, match (yes/no), and confidence score if available. Expect to see: most cells matching cleanly, and a visibly uneven pattern of misses concentrated in specific fields or specific document types, not randomly scattered.
Add three validation-rule columns computed from the extracted values alone: do line items sum to the subtotal (within a cent), is due date on or after invoice date, and does the invoice number match your suppliers' known formats. Flag any row that fails a rule even if every individual field "matched" — this is where the sneaky line-item-swap error from Section 4 tends to surface.
Compute the five metrics from Section 4 from this one spreadsheet: field-level accuracy per column, exact-match rate overall, validation-rule pass rate, confidence calibration (if the confident predictions really were the correct ones), and document-level pass rate (every critical field and every rule passing). Troubleshooting tip: if document-level pass rate is much lower than every individual field-level accuracy, that gap is validation rules catching row-mislabeling and cross-field errors that per-field scoring alone missed — don't discard it as a bug in your scoring, it's the point.
This 20-document exercise is a toy version of the same thing a production pipeline needs at hundreds or thousands of documents and on a recurring cadence — the mechanics don't change, only the volume and the automation around re-running it do.
| Field | Field accuracy | Business weight | Notes |
|---|---|---|---|
| Invoice number | 98% | High | One miss was a PO number substituted in its place. |
| PO number | 85% | High | Missing on documents that never had one — model guessed instead of returning null. |
| Amount due | 90% | Critical | Errors concentrated on poor-scan documents, not clean ones. |
| Tax amount | 80% | Critical | Weakest field; worth its own root-cause pass before automating. |
| Line items (row-level) | 92% | Critical | Validation rule caught 2 additional row-mismatch cases field accuracy missed. |
| Document-level pass | 70% | — | Every critical field and rule passing together — the honest automation-ready rate. |
That last row is the number worth showing to a finance stakeholder: not 90–98% field accuracy, but a 70% document-level pass rate, which is a far more honest answer to "how much of this batch can we trust without a human looking at it."
11. Common Mistakes
- Relying on a single aggregate score. As covered throughout, it averages away exactly the information — which field, which document type, which failure pattern — that decisions actually need.
- Treating confidence as correctness. A confidence score reflects the model's certainty, not ground truth; without calibration checking (does 95% confidence really mean 95% correct on your data), it's a number that feels like safety without necessarily providing it.
- No held-out test set, or test-set leakage. Evaluating on documents the model (or its prompt-engineering process) has effectively already seen inflates every metric and hides real-world performance until production reveals it.
- Ignoring "valid JSON" versus "correct values" as the same thing. Covered in Section 6 — schema compliance and value correctness are different properties, and structured-output features make it easy to assume the first implies the second.
- Treating an offline eval pass as sufficient without live monitoring. A test set is a snapshot; supplier templates change, new currencies appear, scan quality drifts with a new office scanner — without production monitoring, a model that degrades after launch degrades silently.
- Letting the test set go stale. The same test set run for two years without updates increasingly reflects an old version of "what documents look like," not the current one — see the versioning point in Section 8.
- Ignoring cost and latency as evaluation dimensions. An extraction approach that's marginally more accurate but meaningfully slower or more expensive per document can be the wrong production choice even if it wins on accuracy alone — both have to be weighed together against the automation volume.
12. Practical Deployment Checklist
- Define the critical fields for your document type and rank them by financial/business consequence, not extraction difficulty.
- Build a test set covering clean documents, poor scans, rotated pages, tables, handwriting, multiple templates, and missing-field cases, with a version number attached.
- Hand-label ground truth for every critical field, including line items at the row level.
- Compute all five metrics: field-level accuracy, exact-match rate, validation-rule pass rate, confidence calibration, and document-level pass rate.
- Write validation rules independent of the extraction model — sum checks, date-order checks, format checks, cross-references to a vendor or PO master where possible.
- Set confidence thresholds per field, calibrated against your own data, not a vendor's default.
- Define routing rules for automation versus human review, including a random-sample audit of auto-approved documents.
- Gate any prompt, model, or logic change behind the full metric suite before it ships.
- Stand up production dashboards split by pipeline stage, plus alerting on regressions in document-level pass rate or any critical field.
- Schedule a recurring test-set refresh and a recurring review of production failure patterns — this isn't a one-time build.
- Report straight-through processing / document-level pass rate segmented by document type and supplier, not as one blended company-wide figure, and treat any vendor's headline touchless-rate claim as a demo number until it's verified against your own production traffic.
❓ FAQ
Isn't 95% overall accuracy good enough for most document extraction use cases?
It depends entirely on where that 5% falls. If it's spread evenly across low-stakes fields, 95% may genuinely be fine for automation. If it's concentrated on amount due or tax, that same headline number can mean a predictable stream of payment errors — which is why field-level breakdown matters more than the aggregate figure.
Can validation rules replace the need for field-level accuracy metrics?
No — they catch different things. Validation rules catch internal inconsistency (totals that don't add up); field-level accuracy catches whether individual values match ground truth even when they're internally consistent. Both are needed, and neither substitutes for the other.
How often should a document extraction test set be refreshed?
There's no universal cadence, but a good trigger is any of: a new supplier template appearing in volume, a shift in scan quality or intake channel, or simply a defined interval (quarterly is common in practice) so the test set doesn't quietly drift away from production reality.
Should low-confidence fields always be auto-rejected instead of sent to a human?
Usually not — auto-rejecting just moves the bottleneck without resolving it. Routing low-confidence or rule-failing fields to a human reviewer, while still auto-approving the high-confidence majority, is what keeps the system both fast and accurate; the goal is to shrink the review queue over time by fixing root causes, not to avoid review entirely.
Does using an LLM instead of traditional OCR-plus-rules extraction change any of this?
The evaluation principles stay the same — you still need field-level accuracy, validation rules, and staged evaluation — but LLM-based pipelines add one specific new risk worth watching closely: confidently well-formatted, plausible-looking output that's still factually wrong, which is exactly why schema validity and value correctness need to be measured as two separate numbers.
🔗 References & Further Reading
Official/primary documentation consulted for accuracy verification:
- Microsoft — "Interpret and improve accuracy and confidence scores," Azure AI Document Intelligence documentation: learn.microsoft.com
- Amazon Web Services — "Best practices," Amazon Textract Developer Guide: docs.aws.amazon.com
- Amazon Web Services — "Using Amazon Textract with Amazon Augmented AI for processing critical documents," AWS Machine Learning Blog: aws.amazon.com/blogs/machine-learning
All product names (Azure AI Document Intelligence, Amazon Textract, Amazon Augmented AI/A2I) are trademarks of their respective owners.
📝 Summary
- One accuracy score can't tell you which field is failing or how much that failure costs.
- Document AI pipelines have five stages — input, OCR/vision, extraction, validation, human review — and each needs its own check.
- A representative test set deliberately includes poor scans, rotated pages, tables, handwriting, multiple templates, and missing fields.
- Field accuracy, exact-match, validation rules, confidence calibration, and document-level pass rate work as a set, not individually.
- Wrong amounts, PO/invoice-number confusion, and misassigned line items are the highest-consequence, easiest-to-hide failure modes.
- Valid JSON proves the output parses — it says nothing about whether the values are correct.
- Routing decisions should combine confidence and validation results, with a random audit of auto-approved documents.
- Straight-through processing rate is document-level pass rate wearing a business hat — report it segmented, and treat vendor demo numbers skeptically until verified in production.
- Enterprise rollout needs ownership, versioned test sets, CI-gated evaluation, access control, and regression alerting.
- A 20-document scorecard exercise scales directly into a production evaluation pipeline.
If there's one habit worth taking from all of this: before trusting any extraction number, ask "accuracy on which field, on which kind of document?" — that one question is most of the job. Good luck out there. 👋
Comments
Post a Comment