← All articles
AI Engineering6 min read

RAG is not a search problem (and that's why your RAG app is mediocre)

Your RAG app probably isn't failing at search — it's failing at chunking, ranking, and prompting. The fixes that actually move answer quality in production.

RAG is not a search problem (and that's why your RAG app is mediocre)

Teams treat RAG as "search plus a language model," then wonder why answers are vaguely wrong. The search part is usually fine. The failure is everything around it: what got chunked, how it got ranked, and what the model was told to do with the results.

Chunking is the first silent killer. Naive fixed-size chunks split tables, sever procedures mid-step, and orphan headings from their content. We chunk by document structure — sections, tables, list items stay whole — and attach metadata (document type, date, section path) to every chunk. Retrieval then filters on metadata before ranking on similarity, which alone fixed a whole class of "it answered from the wrong document" bugs.

The second killer is the ranking illusion. Vector similarity measures "sounds related," not "answers the question." We run hybrid retrieval — BM25 keyword search plus vectors, fused — because exact terms matter enormously in domains like healthcare and legal, where one wrong word changes the meaning. Then a cross-encoder reranker scores the top candidates properly. Reranking is the single highest-ROI step in most RAG pipelines we've tuned; it regularly beats switching to a bigger embedding model.

Third: the prompt. Most RAG prompts say "answer using the context." That's not enough. Ours say what to do when context is insufficient ("say you don't know"), how to cite (document + section), and what to do with conflicting sources (surface the conflict, don't pick a winner silently). The model needs policy, not just context.

Finally, eval the retriever independently. Build 50 question-answer pairs from your real corpus, measure whether the right chunks come back, and tune chunking and ranking against that — not against vibes. Every RAG rescue project we've done started with this measurement, and every one found the retriever was the problem, not the model.

RAG done right is genuinely powerful — it's the backbone of the clinical knowledge features in HMIS Pro, where a wrong answer isn't just embarrassing, it's dangerous.

Field note: In the clinical knowledge base behind HMIS Pro, metadata filtering does the heavy lifting: every chunk carries document type (protocol, formulary, guideline), effective date, and department. A query about pediatric dosing never even sees adult protocols — they're filtered before ranking. Then the cross-encoder reranker scores the survivors. When we A/B tested reranking versus simply retrieving more chunks, the reranked top-5 beat the unranked top-20 on answer accuracy while cutting context tokens by two-thirds. Better answers, lower latency, lower cost — the rare triple win. If your RAG pipeline has exactly one tuning knob, make it the reranker.

Building something worth shipping?

We take on a small number of AI product engagements. Tell us what you are building — we reply within 48 hours.