← All writing
Retrieval & RAGSeptember 20263 min read

Good retrieval comes before a good answer

A fluent answer can still be built on the wrong paragraph. Making retrieved evidence visible is a useful first step toward understanding a retrieval-augmented system.

On this page

Two separate jobs

Retrieval-augmented generation combines a search step with a generation step. The retriever selects material from an external collection; the generator uses that material to help answer a question. This creates a place to update evidence without treating every document change as a model-training task.

A typical document RAG pipeline. The document index is prepared ahead of time; each question uses that index to retrieve context for an answer.

These two jobs can fail independently. The right passage might never reach the model. Or the passage might be present while the response misreads it. A single answer-quality score cannot tell you which stage needs attention.

Keep the source attached

Start by preserving a document ID, title, section, and location with every chunk. If a search result cannot be traced back to the original text, debugging the answer becomes much harder. A citation needs to lead to evidence a reader can inspect.

Chunk size is an experiment, not a universal constant. A very short chunk can lose the definition that makes a sentence meaningful. A long chunk can mix unrelated subjects. Try paragraph or section boundaries first, then inspect retrieval results for your own questions.

Search the collection

In dense retrieval, an embedding model maps the query and document passages into a shared vector space. A similarity measure ranks candidate passages. Use a model appropriate for short questions matched against longer documents; symmetric sentence similarity and question-to-passage retrieval are different tasks.

Keep a lexical baseline in the comparison, especially for exact names, identifiers, and error messages. If dense search misses a specific function name, changing the generation prompt will not repair the missing evidence.

  1. Write a question with a known answer in the collection.
  2. Inspect the top results and their surrounding text.
  3. Record which relevant passages were found or missed.
  4. Only then inspect the generated answer and its citations.

Evaluate the stages separately

For a small evaluation set, label the passages that support each question. Recall@k asks what fraction of those relevant passages appears among the first k results. Review answer support separately: do the cited passages actually justify the claims in the response?

Include questions that have no answer in the collection. A system that always produces confident prose may look convincing in a demo and still be unhelpful in everyday use. Record missing evidence, ambiguous questions, and retrieval failures as distinct cases.

Sources & further reading

Primary references for the concepts and APIs in this article.

  1. Lewis et al. — Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (opens in a new tab)
  2. Sentence Transformers — Semantic search (opens in a new tab)