All articles
engineering notes · Experiment/prototype

A Computer-Vision Demo Is Not an Accuracy Claim

By SyntaxLab · 2 min read

Counting objects and drawing boxes is only the start; reliable vision work needs representative ground truth and explicit error measures.

A Computer-Vision Demo Is Not an Accuracy Claim

An inspection system can look impressive while drawing boxes on a warehouse photo. The meaningful question is whether those boxes correspond to the objects and defects that matter in the real operating environment.

That requires an evaluation plan before a performance claim.

Define the task precisely

“Inspect the image” is too vague. A measurable version names the target object, the statuses, the image conditions, and the required output. For example: count every carton, flag damaged cartons and missing labels, and return a bounding box for each visible item.

SyntaxLab's inspection demo takes this approach. Its schema defines a scene, an item type, bounding boxes, statuses, defect notes, confidence, and recommendations. The demo supports user uploads and bundled scenes with known ground truth.

Use ground truth, not visual intuition

For sample scenes, the code matches predicted and known boxes with intersection over union (IoU), then reports matched items, false positives, misses, status agreement, and mean overlap. That is an evaluation mechanism, not evidence of a particular accuracy result.

The distinction matters. A published accuracy number needs a documented dataset, labelling process, matching threshold, sampling method, and conditions. It should also state whether it applies to warehouse lighting, camera height, object types, occlusion, and image quality similar to those the customer will use.

Preserve coordinates carefully

Coordinate tasks need careful image handling. The demo reads image dimensions, accounts for certain JPEG orientation cases, asks the model for pixel boxes, then maps them to the browser overlay coordinate system. This avoids pretending that a box from one coordinate space belongs directly in another.

OpenAI's vision documentation notes that coordinate-sensitive tasks need suitable image detail and coordinate mapping. OpenAI Images and Vision guide

Design for review

Inspection outputs should let an operator see the source image, toggle boxes, inspect defects, export findings, and correct mistakes. In high-consequence settings, the result should route to human review rather than directly changing stock, rejecting goods, or issuing penalties.

Project evidence

The evidence is syntaxlab-backend/src/inspection.ts and syntaxlab/src/routes/demos.inspection.tsx. They implement a vision-inspection demo and verification against bundled sample ground truth. No accuracy figure, industrial deployment, or suitability for safety-critical inspection is established by the repository.