All articles
engineering notes · Implemented demo

How to Build a Knowledge-Base Assistant That Knows When It Doesn't Know

By SyntaxLab · 4 min read

Reliable business assistants retrieve evidence, preserve tenant boundaries, and say when the answer is missing.

How to Build a Knowledge-Base Assistant That Knows When It Doesn't Know

A chat interface is not a knowledge system merely because it produces fluent answers. For a business, the important question is whether it can answer from the current policies, prices, products, and procedures it has been given—and clearly say when the evidence is missing.

That is the job of a grounded knowledge-base assistant: retrieve relevant source material before answering, then restrict the answer to that material.

Start with the access boundary

Search must be scoped before it is ranked. In a multi-business product, every retrieval query should be constrained by a tenant, workspace, or visitor identifier in the database—not by a line in a prompt. Otherwise, a good semantic match can still be the wrong customer's data.

SyntaxLab's demo associates uploaded documents and chunks with a visitor-specific identifier derived from a salted hash of an IP address. That supports a no-account demo and avoids storing the raw IP. It also has a stated trade-off: people behind the same network may share a scope. A production product would normally use authenticated business and user identities instead.

Chunk documents for retrieval

Source documents need to be divided into searchable units. Chunks that are too large dilute relevance; chunks that are too small lose the exception or definition that gives a policy meaning. Modest overlap helps preserve meaning at boundaries.

The demo uses paragraph-aware chunks of roughly 1,200 characters, with a 200-character overlap, and includes the filename when creating embeddings. Embeddings turn text into vectors that can be compared for relatedness; semantic search retrieves the material most related to the query. OpenAI's embeddings guide explains this retrieval pattern.

Combine semantic and exact search

Semantic retrieval is good at intent, but exact terms still matter. Product codes, named policies, and prices can be poorly ranked when a search relies only on meaning. The demo first finds cosine-similar chunks, rejects weak matches, then uses PostgreSQL full-text search to add exact keyword hits.

This is a practical hybrid approach. PostgreSQL's websearch_to_tsquery converts a web-style query into full-text-search syntax, while pgvector provides cosine-distance search. PostgreSQL documentation and pgvector's documentation describe the two pieces.

For a small, visitor-scoped document set, the demo deliberately uses exact vector search rather than an approximate index. pgvector notes that approximate indexing trades recall for speed. That trade-off should be measured against the size and latency needs of the actual product, rather than copied from a generic architecture diagram.

Treat follow-ups as a retrieval problem

Users say “What about the premium plan?” rather than repeating the full subject. The demo has a planning step that classifies small talk and rewrites factual follow-ups into standalone search queries. Its role is narrow: improve the query, not introduce new facts.

Make “I couldn't find that” useful

Grounding only helps when the fallback is honest. For factual questions, the demo directs the assistant to answer only from retrieved passages. When nothing relevant is found, it says so and prompts the visitor to add or ask for the missing information.

This is a product choice, not a failure state. Show the source files used, flag low-confidence retrieval, let owners correct stale material, and keep an unanswered-question queue. Those signals are far more useful than a confident answer with no provenance.

Reuse retrieval for voice

Voice should not become a separate, less controlled intelligence layer. SyntaxLab's calling demo connects an ElevenLabs agent to the same server-side knowledge-base search used by chat. The server handles the lookup and call quota; the browser does not receive the upstream API key.

This maps cleanly to the tool-calling pattern: the model requests a tool, the application runs it, and the result is returned for the final response. OpenAI's function-calling guide describes that application-controlled loop.

Grounding is not a guarantee

Retrieved material can be stale, incomplete, ambiguous, or malicious. Document ingestion and retrieved content also need defenses against prompt injection. OWASP identifies prompt injection and vector/embedding weaknesses as important risks for LLM applications. OWASP Top 10 for LLM Applications

The practical test is simple: evaluate real questions, expected sources, failures, and actions before launch. A trustworthy assistant is not one that always sounds certain. It is one whose certainty matches its evidence.

Project evidence

This draft relates to SyntaxLab's implemented text and voice knowledge-base demos: per-visitor ingestion, chunking, embeddings, hybrid retrieval, server-side voice tools, short-lived access tickets, quotas, optional private file storage, and scheduled deletion. It does not claim measured accuracy, customer deployment, or production performance. See syntaxlab-backend/src/kb.ts, chat.ts, db.ts, and src/agents/demo-agent/relay.ts.