Retrieval

Embeddings & RAG Intermediate

Retrieval is the step in RAG that finds and returns the most relevant pieces of text for a given query, either by meaning (semantic search over embeddings), by exact terms (keyword search), or both at once (hybrid search).

In simple terms

It's the part of RAG that answers "which documents should the model actually see." Get this step wrong, missing the right chunk or burying it under irrelevant ones, and no amount of prompting fixes the answer downstream.

Why it matters

Semantic search alone can miss exact matches (a product code, a specific error message) that a keyword search catches instantly, which is why most production systems combine both rather than relying on embeddings alone.

How it works

Semantic retrieval embeds the query and searches a vector database for nearby chunks. Keyword retrieval, commonly BM25, scores documents by term overlap and frequency, closer to how a traditional search engine works. Hybrid search runs both and merges the results, often with a second, more precise reranking pass afterward.

Where it fits

Query → Semantic search + keyword search (hybrid) → Candidate chunks → Reranking → Top-K context

Production impact

Retrieval's own latency, the vector search plus any reranking, sits before the LLM even starts generating, so it directly adds to a RAG system's time to first token.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →