ACUMEN ENGINEERING PERSPECTIVES / AI ENGINEERING

RAG is a retrieval problem before it is a language-model problem

Chunk boundaries, filtering, ranking, and source lifecycle determine whether the model ever sees the evidence it needs.

4 MIN READTECHNICAL APPROACH + WORKED EXAMPLEFOR ENTERPRISE TEAMS

When a RAG answer is wrong, changing the generation prompt may not help. The relevant passage may never have been retrieved. It may have been split away from its table header, filtered incorrectly, or displaced by several semantically similar but irrelevant passages.

We debug retrieval in layers: source coverage, parsing, chunks, embedding behavior, filters, ranking, context assembly, and finally generation. This prevents a larger model from becoming an expensive way to mask a weak index.

RELEVANT TOOLS & TECHNOLOGIES

Selected for the workload—not prescribed as a single mandatory stack. Explore the technology ecosystem ↗

WORKED EXAMPLE / ILLUSTRATIVE, NOT A CLIENT CLAIM

A technical question with an exact part number

A maintenance user asks about an alarm on a specific equipment model. A semantic query retrieves passages about similar alarms across several models, while the correct manual uses a terse identifier that carries little semantic meaning.

Lexical retrieval preserves the model number and alarm code. Semantic retrieval contributes explanatory passages. Their results can be fused, filtered by model and version, and reranked before the answer context is assembled.

The test set should include exact identifiers, ambiguous descriptions, multi-document questions, and queries with no answer. Measure whether the relevant source was retrieved before judging how well the model summarized it.

THE INPUT BOUNDARY

Query + metadata boundary + lexical/semantic candidates

THE USEFUL OUTPUT

Ranked evidence with provenance and source coverage

ACUMEN / ENGINEERING NOTETwo retrieval paths, one evidence setFIG. RAG
Two retrieval paths, one evidence setLexical and vector candidates meet in a filtered, ranked context; the answer layer receives selected evidence rather than the entire repository. Components: Parsed sources; Metadata + versions; Vector index; Exact-term search; Fusion + reranking; Grounded context. ACCESS FILTERS APPLY BEFORE EVIDENCE REACHES THE MODEL 01Parsed sources02Metadata + versions03Vector index04Exact-term search05Fusion + reranking06Grounded context
Lexical and vector candidates meet in a filtered, ranked context; the answer layer receives selected evidence rather than the entire repository.Scroll the drawing sideways to inspect it.

Chunking is an information-design choice

A fixed token window is convenient, but it can separate a definition from its qualification or a table row from its column labels. Preserve headings, source identifiers, page information, and structural context when assembling chunks. Overlap can help continuity while also introducing redundant candidates.

Vector-store selection should consider filtering and operations, not just similarity search. Assess ingestion updates, deleted records, tenant isolation, index build behavior, backups, and how approximate search interacts with the workload. A dedicated vector database and pgvector solve different deployment questions; neither is universally the right choice.

Design the ingestion lifecycle

Connectors, parsing, OCR where needed, chunking, and metadata extraction determine what the index contains. Incremental updates and deletions must carry through to searchable content, with visibility into failed or stale ingestion jobs.

Choose a vector store for the workload

We assess filtering, tenant isolation, index behavior, update patterns, backup, deployment, and operational support. A dedicated vector database or vector capabilities in an existing database can both be appropriate; selection follows the constraints.

Combine semantic and exact retrieval

Embeddings can surface related meaning while lexical search preserves identifiers and exact phrases. Hybrid retrieval and reranking are tested against a baseline. Chunk size and overlap are experiments, not universal settings.

Evaluate each layer separately

A wrong answer may come from missing content, weak retrieval, insufficient context, or generation. We track retrieval relevance, answer support, citation quality, latency, and cost separately so improvements target the actual failure.

Illustrative retrieval experiment—not fixed production settings
query_set → lexical baseline
          → dense retrieval
          → hybrid candidates

compare: relevant-source recall, ranking, permissions
then:    context assembly → answer evaluation

Choose the approach for the constraint

When this mattersAn approach to considerWhat not to assume
Exact identifiers coexist with natural languageHybrid retrieval and evaluated rank fusionMore candidates are not necessarily better context.
Tables or procedures carry meaningStructure-aware parsing and contextual chunksFlat text can lose the relationship needed for the answer.
The system returns confident wrong answersInspect retrieved evidence before changing the modelGeneration evaluation alone cannot locate the failure.

The boundary we keep explicit

RAG reduces reliance on a model’s internal knowledge; it does not eliminate hallucination or guarantee permission safety. Both require explicit tests and controls.

What a useful evaluation should reveal

Evaluate this workload against representative examples and agreed consequences—not just a convincing response. The review should make these dimensions visible:

  • Recall and ranking quality on test queries
  • Groundedness and citation support
  • Index freshness and deletion propagation
  • Latency, storage, and inference cost

Where this approach fits

  • Enterprise RAG and permission-aware search
  • Semantic discovery across technical material
  • Retrieval services embedded in enterprise tools

A considered first step

Create a retrieval benchmark from real questions and reviewed source passages. Compare an exact-search baseline, semantic search, and a hybrid pipeline before optimizing the answer layer.

Serving enterprise teams in California, Atlanta, Georgia, and across the United States.

Discuss your requirements

Questions worth resolving

Which vector database should we use?

The answer depends on workload, deployment, operational skills, filtering needs, and existing data infrastructure. We make the choice through a requirements review and a focused benchmark.

Can retrieval respect document permissions?

Yes, but permissions must be designed into indexing and query execution. We test user and tenant isolation instead of assuming the model will refuse unauthorized information.

A CONVERSATION IS A GOOD START

Let’s put your
ideas to work.

Choose a time to talk, or leave your email and a little context. We’ll take it from there.

Book a 15-minute call
OR LET US GET IN TOUCH
Prefer your email app? business@acumen.llc ↗