Fixing RAG Retrieval Failures: The "Smart" Stack
You built a RAG (Retrieval Augmented Generation) system. It works great for questions like "What is our vacation policy?" but fails miserably on specific queries like "What is the error code in log #9921?"
The problem isn't the LLM; it's the Retrieval. Standard vector search is often too "fuzzy" for precise facts. This guide covers the three industry-standard upgrades to fix this.
1. Hybrid Search (The Best of Both Worlds)
Why Vector Search Fails
Vector databases search for concepts. If you search for "Apple", it might return "Fruit" or "Technology". But if you need an exact match for a part number like XJ-9000-Z, vector search often struggles because it tries to find the "meaning" of a serial number.
The Fix: Add Keywords (BM25)
Hybrid Search runs two queries simultaneously:
- Sparse Vector (Keyword/BM25): Looks for exact word matches.
- Dense Vector (Semantic): Looks for similar meanings.
It then merges the results using a technique called Reciprocal Rank Fusion (RRF) to give you the best candidates.
2. Reranking (The Quality Control)
Retrievers (like Pinecone or Chroma) optimize for speed, not perfection. They might return 100 documents, but the "perfect" answer is often buried at position #45, where the LLM might miss it.
How Reranking Works
You implement a Two-Stage Retrieval process:
- Retrieve: Get the top 50 loosely relevant documents using fast Hybrid Search.
- Rerank: Use a specialized "Cross-Encoder" model (like Cohere Rerank or BGE) to score every document against the query deeply.
- Select: Keep only the top 5 highest-scored chunks for the LLM.
3. Knowledge Graphs (GraphRAG)
Sometimes the answer isn't in one document—it's scattered across many.
The Problem
"How is the CEO of Company A related to the founder of Company B?"
Vector search fails here because no single document contains the entire path.
The Graph Solution
Knowledge Graphs store data as nodes and edges (relationships).
(Person A) --[worked_at]-- (Company A) --[partnered_with]-- (Company B)
GraphRAG combines vector search with graph traversal. It can "hop" from node to node to answer complex multi-step reasoning questions that flat text retrieval simply cannot solve.
Putting It All Together (Code)
Here is a conceptual Python example using LangChain and a hypothetical Reranker to demonstrate the flow.
from langchain.retrievers import EnsembleRetriever
from langchain_community.retrievers import BM25Retriever
from langchain_core.documents import Document
# 1. Setup Base Retrievers
# BM25 for Keyword Match (Sparse)
bm25_retriever = BM25Retriever.from_documents(docs)
bm25_retriever.k = 10
# Vector for Semantic Match (Dense)
vector_retriever = vectorstore.as_retriever(search_kwargs={"k": 10})
# 2. Hybrid Search (Ensemble)
# combining results 50/50
ensemble_retriever = EnsembleRetriever(
retrievers=[bm25_retriever, vector_retriever],
weights=[0.5, 0.5]
)
# 3. The Reranking Step
def rigorous_retrieval(query):
# Step A: Get top 20 candidates (Hybrid)
initial_docs = ensemble_retriever.invoke(query)
# Step B: Rerank (using a Cross-Encoder like Cohere)
# Ideally, this uses a specialized model API
ranked_docs = reranker_client.rerank(
query=query,
documents=initial_docs,
top_n=5
)
return ranked_docs
# Usage
results = rigorous_retrieval("Error code 503 in production logs")
# Result: High precision documents only.Conclusion
Fixing retrieval failures is about moving from "naive" single-step search to a robust multi-stage pipeline.
- Use Hybrid Search to catch exact keywords.
- Use Reranking to filter noise and fix ordering.
- Use Knowledge Graphs when relationships matter more than similarity.
Start with Hybrid and Reranking—they are the easiest wins. Save GraphRAG for when you have highly interconnected data.