
The hardest part of a business workflow is often understanding an unstructured input: an invoice, email, contract, or customer request. Large language models are useful here. They should not, however, become the final authority for arithmetic, access control, payment, or approval policy.
Separate interpretation from action
A dependable workflow has a clear handoff:
- A model classifies and extracts facts from the unstructured input.
- The application validates those facts and checks known records.
- Deterministic rules decide the next permitted step.
- A human approves high-impact actions.
- The system records what happened and why.
SyntaxLab's invoice-workflow demo makes this separation explicit. AI reads an invoice, while code checks a sample vendor master, purchase order, banking details, totals, approval thresholds, and the workflow audit trail. The demo data is fictional, so it should be described as an architectural demonstration—not an accounts-payable integration.
Why deterministic rules matter
If an amount must equal the sum of line items, calculate it. If an invoice above a threshold needs finance approval, express that threshold in code or a versioned policy. If a payment can be scheduled only after an approval, enforce that state transition in the application.
This produces an explanation a finance or operations team can audit: which document fields were extracted, which check passed or failed, who approved, and what changed next.
Keep tools narrow and permissioned
Tool calling can connect models to application data and actions, but a model's request is not authorisation. Validate arguments, check permissions, make side-effecting operations idempotent, and require approval for high-impact actions. OpenAI's tool-calling guidance specifically recommends application-side argument and permission checks, idempotency where possible, and approval before high-impact actions. OpenAI tool-calling guidance
Design for exceptions, not only happy paths
Real workflows need states for missing vendor records, conflicting totals, duplicate invoices, downstream service outages, and rejected approvals. These should be first-class statuses, not generic error messages that leave users to guess what happened.
The practical measure of quality is not how impressive the first AI extraction looks. It is whether the system reaches a safe, explainable outcome when information is incomplete or wrong.
Project evidence
The source for this draft is syntaxlab-backend/src/workflow.ts, which implements the invoice workflow and audit model, plus src/inbox.ts, which uses structured AI analysis but calculates catalogue pricing, discounts, and VAT in code. Both are surfaced as demos. No customer accounting system, payment rail, or financial outcome is evidenced in this repository.