
An invoice-processing system should not stop at “the model found an amount.” A number that looks plausible can still be a wrong total, a missing tax registration number, or a mismatch between line items and the printed balance.
The practical architecture is two-part: let AI read the messy document, then let ordinary software check the things ordinary software can prove.
Use AI for interpretation
Business documents vary widely. A receipt photographed on a phone, a supplier PDF, a bilingual quotation, and a scanned purchase order rarely share one template. This is where a vision-capable model is useful: identify the document type, issuer, recipient, references, dates, line items, totals, payment details, and visible issues.
SyntaxLab's document-intelligence demo defines a schema for these fields, including explicit null values when information cannot be read. Its interface streams partial structured results so a person can see work progress rather than wait for an opaque final screen.
Structured output is valuable because downstream software needs known fields and types, not a paragraph that happens to mention a total. OpenAI documents Structured Outputs as a way to constrain model output to a supplied JSON schema, while still requiring applications to handle refusals and incomplete responses. OpenAI Structured Outputs
Keep maths and policy outside the model
Once the model returns fields, code should calculate the arithmetic again: sum line items, apply stated discounts, compare tax amount and total, and detect discrepancies. Similarly, a policy rule should check required identifiers or approval thresholds in deterministic code.
This separation is not anti-AI. It assigns work to the component best suited to it:
- The model interprets layouts and language.
- A schema gives the application a predictable handoff.
- Code performs calculations and repeatable policy checks.
- A person resolves exceptions and approves consequential action.
For file inputs, capabilities vary by file type. OpenAI documents that vision-capable processing of PDFs can include both extracted text and page images, while many non-PDF office formats are processed as text. OpenAI file inputs guide This is why teams should test the actual document mix they expect, not only clean sample PDFs.
Design an exception queue
The right output is not always “processed.” A useful system identifies what needs attention: low legibility, missing values, arithmetic differences, unusual suppliers, or documents that do not fit the expected category.
That means preserving the original document, extracted data, validation findings, and an audit record of changes. It also means showing a reviewer the specific source material behind a field, not only a generated summary.
What to test before automating
Build an evaluation set from representative documents: clean digital files, scans, phone photos, Arabic/English combinations, long invoices, credit notes, and deliberately flawed examples. Measure field accuracy and validation outcomes separately. A system can extract a vendor name well while still failing at table boundaries or tax math.
Project evidence
This article is grounded in the SyntaxLab document-intelligence demo. syntaxlab-backend/src/extraction.ts defines structured extraction, accepts PDFs, images, and text documents, and performs server-side arithmetic and selected checks. The UI lives in syntaxlab/src/routes/demos.documents.tsx. This is prototype/demo evidence only; it does not establish regulatory compliance, extraction accuracy, or an ERP deployment.