OCR is one of those AI capabilities that looks solved until you put it near money, onboarding, compliance, claims, or procurement. The demo is easy: upload an invoice, extract the vendor, date, total, and line items, then send the JSON to the next system.
That is the weak shape. A document extraction system is not production-ready because it produced JSON. It becomes production-ready when the system can say which parts of the JSON are supported by visual evidence, which fields were low confidence, which totals do not reconcile, and which records must stop before they become business facts.
In this issue, we build a .NET 10 OCR extraction contract for invoice processing. Azure Document Intelligence is the live extraction provider, because its layout and invoice models return structured fields, tables, confidence values, and bounding regions. The companion repo runs live against Azure by default, while the unit tests use in-memory extraction objects to protect the deterministic contract. Azure extracts the document; deterministic code decides whether that extraction is safe enough to enter the workflow.
The First Mistake Is Trusting The OCR
When we say OCR, we often collapse several different jobs into one word. There is raw text recognition. There is layout detection. There is table reconstruction. There is key-value extraction. There is document-type knowledge, like knowing that an invoice probably has a vendor, an invoice ID, dates, tax, totals, and line items.
Those jobs have different failure modes. A number can be read correctly but assigned to the wrong field. A table row can be detected with weak confidence. A total can be extracted cleanly even though the line items do not add up. A field can have a value but no useful bounding region, which means a reviewer cannot easily verify where the system found it.
Azure AI Document Intelligence gives us a strong provider boundary for this problem. The layout model can extract text, tables, selection marks, and bounding regions. The invoice model adds invoice-specific fields and line items. That is useful, but it does not remove the application contract. It gives us better evidence to validate. The screenshot below shows that boundary in practice: extracted fields, values, and confidence scores before our code decides what to trust.

Source: Microsoft Learn, Azure AI Document Intelligence invoice model documentation.
What We Build
The repository is a document intake gate for accounts payable. It reads a manifest of invoice documents, extracts fields through Azure Document Intelligence, normalizes the provider response into an internal contract, and routes each document to one of three states: admitted, needs_review, or quarantined.
That route is not chosen by the OCR provider. It is chosen by code. The provider can say, "I think this field is the invoice total, and here is my confidence." The workflow still decides whether that confidence is good enough, whether the field has visual evidence, whether the currency matches the manifest, and whether the total reconciles with the table.
I like this example because it is ordinary. Invoices are not exotic AI. They are exactly the kind of document workflow where teams are tempted to move quickly, wire the extraction output into a database, and discover the contract problem later during reconciliation or audit.
Azure Extracts, Code Decides
The live extractor calls Azure Document Intelligence with the prebuilt-invoice model. The project uses a small REST adapter instead of hiding the boundary behind a framework. It posts the PDF or image, polls the operation, and maps the Azure result into one internal DocumentExtraction object.
public sealed class DocumentExtraction
{
public string DocumentId { get; set; } = "";
public string Provider { get; set; } = "";
public string ModelId { get; set; } = "";
public int PageCount { get; set; }
public List<ExtractedField> Fields { get; set; } = [];
public List<ExtractedLineItem> LineItems { get; set; } = [];
}That internal contract is intentionally boring. It has fields, line items, confidence scores, and bounding regions. We do not want the rest of the workflow coupled to every detail of the Azure response. We want a stable application shape that can be tested deterministically, while the runtime path still uses the real Azure extraction service.
The app runs live against Azure. The only secret it needs is the Document Intelligence key:
$env:OCRGATE_Extraction__AzureEndpoint = "https://ocr-scanner-modern-engineer.cognitiveservices.azure.com"
$env:OCRGATE_Extraction__AzureModelId = "prebuilt-invoice"
$env:OCRGATE_AZURE_DOCUMENT_INTELLIGENCE_KEY = "<your-key>"
dotnet run --project OcrExtractionContractsThe live adapter also handles the boring operational part: Azure can throttle small free-tier resources during polling. The project retries transient provider failures with backoff, and if retry exhaustion still happens, the document becomes a quarantined extraction_failed record instead of disappearing from the run. The tests still avoid cloud calls. That is deliberate. We want the production app to use the live provider, but we do not want unit tests to become network tests. The gate can be tested with in-memory extraction objects because the contract boundary is small and explicit.
Confidence Is Not A Decoration
A confidence score is only useful if it changes the behavior of the system. In this project, required fields are configured explicitly: vendor name, invoice ID, invoice date, invoice total, and currency. Each one must be present, above the required confidence threshold, and backed by a bounding region. The diagram below is the core decision: a field can be present and visually grounded, but still fail the workflow when confidence sits below the contract threshold.
"Contract": {
"RequiredFields": [
"vendor_name",
"invoice_id",
"invoice_date",
"invoice_total",
"currency"
],
"MinRequiredFieldConfidence": 0.80,
"MinLineItemConfidence": 0.70,
"MaxTotalTolerance": 0.05,
"RequireBoundingRegions": true
}This is where the deterministic boundary starts to matter. If invoice_id is missing, the document is quarantined. If invoice_total is present but below threshold, it is quarantined. If a required field has no bounding region, it is quarantined. We do not ask a language model to infer the missing invoice number from surrounding prose. Missing evidence is a workflow state, not a prompt challenge.
The Table Has To Agree With The Header
Invoice extraction is not only a header-field problem. The table matters. A system that extracts the total but ignores weak line items is not ready for downstream automation. The total may be right, but the cost allocation, purchase order matching, or approval category may still be wrong.
The sample gate treats low-confidence line items as reviewable, not automatically fatal. That is a practical distinction. A weak row might still be fixable by a human reviewer, while a missing required field or contradictory total should stop the document from entering the workflow.
if (field.Confidence < contract.MinRequiredFieldConfidence)
{
reasons.Add($"weak_required_field:{requiredField}:{field.Confidence:F2}");
}
if (contract.RequireBoundingRegions && field.BoundingRegions.Count == 0)
{
reasons.Add($"missing_bounding_region:{requiredField}");
}The total check is stricter. If subtotal plus tax does not match the extracted invoice total within tolerance, the document is quarantined. This isn't an OCR problem. If a document contains conflicting financial evidence, it shouldn't become payable until someone resolves the discrepancy.
Quarantine Is Part Of The Product
A lot of automation projects treat quarantine as a sad path. I think that is backwards. Quarantine is how the system tells the truth. It says: this document might be readable, but it is not strong enough to become a business event.
The repo writes quarantined decisions into a separate JSONL file. That makes the operational conversation concrete. A reviewer can see whether the problem was a missing field, weak confidence, currency mismatch, missing evidence region, total mismatch, or live provider failure. The system is not hiding uncertainty inside a polite generated explanation.
A Live Run Tells The Story
The default command sends the manifest documents to Azure:
$env:OCRGATE_AZURE_DOCUMENT_INTELLIGENCE_KEY = "<your-key>"
dotnet run --project OcrExtractionContractsA clean unthrottled run produces admitted documents and quarantined documents:
OCR Extraction Contracts for Document AI Systems
Provider: azure
Azure model: prebuilt-invoice
Output directory: D:\...\OcrExtractionContracts\data\output
Running live Azure Document Intelligence extraction...
Documents: total=7 admitted=4 needs_review=0 quarantined=3
AP-1001 | ADMITTED | vendor=Northwind Traders | total=248.64 | reasons=clear
AP-1002 | QUARANTINED | vendor=Contoso Office Supply | total=189.50 | reasons=weak_required_field:vendor_name:0.75,missing_required_field:invoice_id
AP-1003 | QUARANTINED | vendor=Fabrikam Cloud Services | total=610.00 | reasons=total_mismatch:expected=540.00:actual=610.00:difference=70.00
AP-1004 | QUARANTINED | vendor=ADVENTURE WORKS | total=340.20 | reasons=weak_required_field:vendor_name:0.72
AP-1005 | ADMITTED | vendor=Kane-Morgan | total=96.73 | reasons=clear
AP-1006 | ADMITTED | vendor=Gutierrez, Shah and Davis | total=25.52 | reasons=clear
AP-1007 | ADMITTED | vendor=Levy-Vargas | total=205.50 | reasons=clearThose lines are the issue. Several invoices are clean enough to admit. One supplier statement has no formal invoice ID and a weak vendor field. One has confident-looking fields but the money does not reconcile. One has the right values, but the vendor evidence is below the required threshold. The system does not need drama. It needs a route.
The Tests Protect The Boundary
The unit tests run without Azure credentials:
dotnet test OcrExtractionContracts.slnxThey cover the cases that would hurt in production:
- clean invoice admission
- missing required field quarantine
- low required-field confidence quarantine
- total mismatch quarantine
- low-confidence line item review
- pipeline report and output file generation
- live provider failure quarantine
We are not testing whether Azure can read every invoice format in the world. We are testing whether our workflow refuses to treat uncertain extraction as confirmed business data. That is the contract we own.
Where I Would Use This
I would use this pattern anywhere a document becomes an operational event: invoice intake, purchase orders, receipts, insurance claims, onboarding forms, tax forms, proof-of-delivery scans, compliance attestations, and vendor packets. The details change, but the control question is the same: what evidence is strong enough to move forward?
I would not make an LLM the first extraction authority here. A vision-capable model can help with downstream review, explanation, or exception handling, but the first boundary should return inspectable evidence: fields, coordinates, confidence, tables, and source pages. After that, code can decide what the model output is allowed to do.
What I Would Add Next
The next version should add a real human review queue. Reviewers should correct missing fields, approve quarantined documents, and mark false positives. Those corrections should become a governed evaluation set, not a pile of unreviewed feedback that quietly changes future behavior.
I would also add source-file checksums, duplicate detection, vendor matching, purchase-order lookup, and per-vendor thresholds. Some vendors send clean digital invoices every time. Others send scans, photos, and layout changes. The contract should eventually learn where each vendor is reliable and where it needs stricter gates.
Finally, I would track extraction drift. If a provider model upgrade suddenly lowers confidence, changes field names, or alters table shape, the workflow should catch that before invoices start flowing into downstream systems with a different meaning.
Final Notes
OCR is useful because it turns static documents into machine-readable evidence. It is dangerous when we forget the evidence part and treat the first extracted JSON as fact.
The production shape is not complicated. Use a real extraction provider. Normalize the result. Declare required fields. Require confidence and visual evidence. Reconcile totals. Quarantine contradictions. Test the gate with deterministic in-memory records. That is enough to move document AI from demo territory into engineering territory.
Explore the repository at the GitHub repository.
See you in the next issue.
Stay curious.
