ProxyAI
Cover image for Grounding a support bot in your own documents, what actually makes retrieval work

Grounding a support bot in your own documents, what actually makes retrieval work

Uploading a PDF is the easy part. What decides whether a grounded assistant answers correctly is chunking, hybrid retrieval, and ruthless discipline about which documents are allowed in.

·7 min read

In short: grounding a support bot means retrieving passages from your own documents at question time and answering from those, instead of from the model's training. Three things decide whether it works: whether retrieval matches both meaning and exact terms, whether source documents are current and non-contradictory, and whether each bot's index is isolated. Uploading more documents makes a bot worse at least as often as it makes it better.

What grounding actually is

When a customer asks a question, the system searches your indexed documents, pulls the passages that match, and gives the model those passages to answer from. The model is not remembering your return policy, it is reading it, at the moment of the question.

That distinction matters for two reasons. Your policy can change on Tuesday and be correct in answers on Tuesday. And when the answer is wrong, the cause is almost always retrieval, the right passage was not found, or the wrong one was, rather than the model.

Semantic search alone is not enough

Most retrieval systems match on meaning: the question and the documents are turned into vectors, and the closest passages win. This handles paraphrase well. Someone asking "can I send this back" finds a passage about returns without sharing a single word with it.

It handles exact strings badly. Part number RX-4410-B, order prefix INV-, a specific error code, these are precisely the queries where semantic similarity is least useful, because the token that matters carries almost no meaning to embed.

The fix is to index everything twice: once semantically, once as keyword terms. A question is matched both ways and the result sets are merged. Paraphrased questions and exact product codes both land. ProxyAI does this by default, which is why a question containing a SKU behaves as well as one containing a sentence.

If you are evaluating a platform, this is the specific thing worth asking about. "We use RAG" says nothing about whether a customer quoting a part number will get an answer.

Chunking decides more than it looks like it does

Documents are split into passages before indexing, and retrieval returns passages rather than whole files. So the split determines what the model can see.

Split too small and a passage arrives without the context that made it meaningful, a price with no indication of which product it belongs to. Split too large and the model is handed three pages to find one sentence in, which dilutes attention and costs more.

The practical lever most teams have is not a chunking setting. It is document structure. A source file with clear headings splits along meaningful boundaries by itself. A wall of undifferentiated prose does not, no matter how the splitter is configured. Rewriting your top ten support documents with real headings is usually a bigger accuracy win than any retrieval tuning.

The discipline that matters most: what you do not upload

The instinct is to upload everything. It is the wrong instinct.

Retrieval returns what best matches the question, not what is most current. Upload last year's shipping policy alongside this year's and a customer will eventually be told the wrong thing, confidently, with the outdated passage as the reason. The bot did not hallucinate, it read a document you gave it.

Three rules that prevent most of this:

One current version of each thing. Delete superseded documents rather than keeping them "for reference". The index has no concept of reference.

Filenames that mean something. On platforms where a same-named upload replaces the previous version, consistent naming is how you update a policy rather than accumulate two contradictory copies of it.

Review before you implant. Extraction is imperfect, especially from PDFs. A table that turned into a wall of numbers during extraction is worse than no table, and you can only see that by looking.

Scanned PDFs need a decision, not a default

A PDF with no text layer, a scan, a photographed page, extracts to nothing at all. OCR fixes this, and adds processing cost and a new failure mode, since OCR output is never perfect.

So it is a per-document decision, not a global setting. Turn it on for documents that genuinely have no text layer; leave it off for the ones that do. ProxyAI runs OCR in the browser, on your machine, so scanned documents are not sent to a third-party OCR service, worth knowing if the scans are contracts.

Isolation is a correctness requirement, not just security

If you run more than one bot, each needs its own index namespace, and a search must be scoped to that namespace.

The obvious reason is confidentiality: one client's bot must not retrieve another client's documents. The less obvious reason is accuracy. Two brands with different return windows in one shared index means retrieval picks whichever passage scores higher, and it will sometimes pick the wrong company's policy. Isolation is what makes the answer deterministic.

How to tell whether it is working

Test in both directions, and the second one is the one people skip.

Ask what is covered. Questions your documents answer directly. Check the answer matches the source rather than a plausible paraphrase of it.

Ask what is not covered. Questions near your documents but not in them. The correct behaviour is admitting it does not know. A bot that confidently improvises here will do the same thing to a customer, and you will not find out until they act on it.

Then check retrieval logs for the queries that returned nothing useful. Those are your content gaps, stated as questions real people asked, a better list of what to write next than anything you would produce in a planning meeting.

FAQ

What does it mean to ground a chatbot in your own documents?

Grounding means the bot searches your indexed documents when a question arrives and answers from the retrieved passages, rather than from the model's training data. The source text is read at question time, so updating a document updates the answers immediately.

What is hybrid search and why does it matter for support bots?

Hybrid search indexes content twice, once semantically and once as keyword terms, then merges both result sets for each query. It matters because semantic search alone handles paraphrased questions well but performs poorly on exact strings such as SKUs, order prefixes and error codes, which are common in support.

Why does my grounded chatbot give outdated answers?

Almost always because an outdated document is still in the index. Retrieval returns the closest match to the question, not the most recent document, so a superseded policy will be quoted whenever it scores higher. Deleting old versions fixes it; keeping them "for reference" does not work, because the index has no concept of reference.

How many documents should I upload to a support chatbot?

Fewer than instinct suggests. Every added document is another candidate for retrieval, so duplicates and superseded versions actively reduce accuracy. One current version of each topic, with clear headings, outperforms a large unmanaged corpus.

Do scanned PDFs work in a chatbot knowledge base?

Only with OCR, since a scan has no text layer to extract. Enable OCR per document rather than globally, because it adds processing cost and imperfect extraction. ProxyAI performs OCR in the browser on your own machine, so scanned files are not sent to an external OCR service.

Can one chatbot retrieve another chatbot's documents?

Not when each bot has its own index namespace and searches are scoped to it. Isolation matters for confidentiality and for accuracy, two brands with different policies sharing one index means retrieval can return the wrong company's answer.

How do I test whether a grounded chatbot is working?

Ask questions your documents cover and check the answer against the source, then ask questions near your documents but not in them and confirm the bot says it does not know. The second test is the one that predicts real-world failures.


The mechanics, supported file types, size limits, OCR and storage billing, are in the Bot Knowledge Guide.

rag · knowledge-base · retrieval

More from the blog