All articles
engineering notes · Practical business guide

A RAG Evaluation Scorecard for Business Knowledge Assistants

By SyntaxLab · 2 min read

Evaluate retrieval and generated answers separately with answerable questions, missing evidence, permission tests, and source-grounded acceptance criteria.

A RAG Evaluation Scorecard for Business Knowledge Assistants

A knowledge assistant can fail because the correct source was never retrieved or because the answer misused a good source. A useful scorecard separates these stages. Otherwise a model change may appear to fix search when it merely produces more convincing text.

A practical scenario

For each test question, record the expected document, supporting passage, permitted user scope, and acceptable answer. Include questions whose answer does not exist in the collection. These cases test whether the assistant can state that the available material is insufficient.

Design the first version

Label retrieval coverage, answer relevance, factual support, citation correctness, and refusal behavior independently. Use human review for ambiguous policy interpretation. Automated grading can help prioritize review but should be checked against human labels before being used as the only release gate.

What to test and measure

Run the same cases after changes to documents, prompts, chunking, filters, or models. Keep a separate set for newly discovered failures. Review results by question category and permission scope, not just one average. A system is ready to expand when its important failure types are visible and its fallback is useful.

Questions to resolve before commissioning

  • Which source system owns the facts used in this workflow?
  • Who reviews exceptions and corrects inaccurate output?
  • What baseline and acceptance criteria will determine whether the pilot is useful?
  • What should the user do when a source, tool, or device is unavailable?

Explore the implementation

This is a planning guide, not a report of measured client results. Examples are illustrative. Explore the related SyntaxLab demo to discuss the interaction, then use your own records and acceptance criteria for a production pilot. Discuss a scoped project or review our AI automation services.

Further reading

Read Microsoft guidance on evaluating RAG answers for technical background. Continue with Multi-Tenant RAG: Protect Business Documents Before Retrieval.