When a RAG answer is wrong, changing the generation prompt may not help. The relevant passage may never have been retrieved. It may have been split away from its table header, filtered incorrectly, or displaced by several semantically similar but irrelevant passages.
We debug retrieval in layers: source coverage, parsing, chunks, embedding behavior, filters, ranking, context assembly, and finally generation. This prevents a larger model from becoming an expensive way to mask a weak index.
Selected for the workload—not prescribed as a single mandatory stack. Explore the technology ecosystem ↗
A technical question with an exact part number
A maintenance user asks about an alarm on a specific equipment model. A semantic query retrieves passages about similar alarms across several models, while the correct manual uses a terse identifier that carries little semantic meaning.
Lexical retrieval preserves the model number and alarm code. Semantic retrieval contributes explanatory passages. Their results can be fused, filtered by model and version, and reranked before the answer context is assembled.
The test set should include exact identifiers, ambiguous descriptions, multi-document questions, and queries with no answer. Measure whether the relevant source was retrieved before judging how well the model summarized it.
Query + metadata boundary + lexical/semantic candidates
Ranked evidence with provenance and source coverage
Chunking is an information-design choice
A fixed token window is convenient, but it can separate a definition from its qualification or a table row from its column labels. Preserve headings, source identifiers, page information, and structural context when assembling chunks. Overlap can help continuity while also introducing redundant candidates.
Vector-store selection should consider filtering and operations, not just similarity search. Assess ingestion updates, deleted records, tenant isolation, index build behavior, backups, and how approximate search interacts with the workload. A dedicated vector database and pgvector solve different deployment questions; neither is universally the right choice.
Design the ingestion lifecycle
Connectors, parsing, OCR where needed, chunking, and metadata extraction determine what the index contains. Incremental updates and deletions must carry through to searchable content, with visibility into failed or stale ingestion jobs.
Choose a vector store for the workload
We assess filtering, tenant isolation, index behavior, update patterns, backup, deployment, and operational support. A dedicated vector database or vector capabilities in an existing database can both be appropriate; selection follows the constraints.
Combine semantic and exact retrieval
Embeddings can surface related meaning while lexical search preserves identifiers and exact phrases. Hybrid retrieval and reranking are tested against a baseline. Chunk size and overlap are experiments, not universal settings.
Evaluate each layer separately
A wrong answer may come from missing content, weak retrieval, insufficient context, or generation. We track retrieval relevance, answer support, citation quality, latency, and cost separately so improvements target the actual failure.
query_set → lexical baseline
→ dense retrieval
→ hybrid candidates
compare: relevant-source recall, ranking, permissions
then: context assembly → answer evaluationChoose the approach for the constraint
| When this matters | An approach to consider | What not to assume |
|---|---|---|
| Exact identifiers coexist with natural language | Hybrid retrieval and evaluated rank fusion | More candidates are not necessarily better context. |
| Tables or procedures carry meaning | Structure-aware parsing and contextual chunks | Flat text can lose the relationship needed for the answer. |
| The system returns confident wrong answers | Inspect retrieved evidence before changing the model | Generation evaluation alone cannot locate the failure. |
The boundary we keep explicit
RAG reduces reliance on a model’s internal knowledge; it does not eliminate hallucination or guarantee permission safety. Both require explicit tests and controls.
What a useful evaluation should reveal
Evaluate this workload against representative examples and agreed consequences—not just a convincing response. The review should make these dimensions visible:
- Recall and ranking quality on test queries
- Groundedness and citation support
- Index freshness and deletion propagation
- Latency, storage, and inference cost
Where this approach fits
- Enterprise RAG and permission-aware search
- Semantic discovery across technical material
- Retrieval services embedded in enterprise tools
A considered first step
Create a retrieval benchmark from real questions and reviewed source passages. Compare an exact-search baseline, semantic search, and a hybrid pipeline before optimizing the answer layer.
Serving enterprise teams in California, Atlanta, Georgia, and across the United States.
Discuss your requirementsQuestions worth resolving
Which vector database should we use?
The answer depends on workload, deployment, operational skills, filtering needs, and existing data infrastructure. We make the choice through a requirements review and a focused benchmark.
Can retrieval respect document permissions?
Yes, but permissions must be designed into indexing and query execution. We test user and tenant isolation instead of assuming the model will refuse unauthorized information.