Shipping a RAG system that doesn't hallucinate
Retrieval quality — not the model — is where RAG lives or dies. A practical guide to chunking, grounding, and citation strategies for production.
The first RAG demo always looks magical. You upload a PDF, ask a question, and the model answers. Then you ship it, a real user asks something slightly outside the documents, and the model confidently invents an answer. That gap — between the demo and the deploy — is almost never about the model. It's about retrieval.
I built a document Q&A system that answers questions across large libraries with source citations, and a brand-visibility platform that runs 100K+ AI queries a day. In both, the same lesson held: if you fix retrieval and grounding, hallucinations mostly disappear. Here's the playbook I use.
1. Chunk for meaning, not for tokens
The default "split every 1,000 characters" destroys the thing retrieval depends on: semantic coherence. A chunk that ends mid-sentence, or merges two unrelated sections, produces a fuzzy embedding that matches everything and nothing. Split on structure first (headings, paragraphs), then pack up to a target size, and add a small overlap so context isn't severed at the boundary.
// Structure-aware chunking beats fixed-size splitting
const chunks = splitByHeadings(doc)
.flatMap((section) =>
packParagraphs(section, {
targetTokens: 400,
overlapTokens: 60, // keep context across the seam
})
)
.map((c) => ({ ...c, sourceId: doc.id, heading: c.heading }));2. Retrieve more, then re-rank
Vector search is fast but approximate. Pull a wider candidate set (top 20–30) with the embedding index, then re-rank that shortlist with a cross-encoder or the LLM itself against the actual query. You get the recall of a big search with the precision of a careful read — without paying to re-rank the whole corpus.
3. Ground the prompt — and let it refuse
This is the single highest-leverage change. Instruct the model to answer only from the retrieved context, to cite the chunk it used, and — crucially — to say "I don't know" when the context doesn't contain the answer. A model that is allowed to refuse stops hallucinating to fill silence.
You are answering strictly from the CONTEXT below.
- If the answer is not in the context, reply exactly:
"I couldn't find that in the provided documents."
- Every claim must cite its source id like [#3].
- Never use outside knowledge.
CONTEXT:
{{retrieved_chunks}}A model that is allowed to say "I don't know" stops inventing answers to fill the silence.
4. Make citations first-class
Return the source chunks alongside the answer and render them in the UI. Two things happen: users trust answers they can verify, and you get a debugging surface. When an answer is wrong, you can immediately see whether retrieval surfaced the wrong chunk (a retrieval bug) or the right chunk was ignored (a prompting bug).
5. Evaluate retrieval separately from generation
Before blaming the model, measure whether the correct chunk was even in the retrieved set. Keep a small golden set of question → expected-source pairs and track hit-rate on every change to chunking or embeddings.
- Retrieval hit-rate: was the answer-bearing chunk retrieved at all?
- Grounding: does every sentence trace to a cited chunk?
- Refusal accuracy: does it decline when it should?
The takeaway
Teams reach for a bigger model when RAG feels unreliable. Nine times out of ten the fix is upstream: better chunks, wider retrieval with re-ranking, a grounded prompt that's allowed to refuse, and citations you can inspect. Get retrieval right and the model has an easy job.
Building something with AI?
I'm Muhammad Usman Haider, a full stack ai engineer. If you want a system like the ones in this article, let's talk.
Start a project