RAG is a pattern where, before an LLM answers a question, the system first retrieves relevant documents from an external source and includes them in the prompt, so the model answers from that specific material instead of only from what it memorized during training.
Instead of asking a model to answer purely from memory, RAG hands it an open book first: the most relevant pages from your own documents, pulled in right before it writes the answer. The model still writes the final response, but it's now grounded in text it was just shown, not just what it learned months or years ago.
It's the standard way to make an LLM answer accurately about information that didn't exist (or wasn't public) when the model was trained, private company documents, yesterday's news, a product manual, without retraining the model itself.
A document collection is chunked and embedded ahead of time and stored in a vector database. At query time, the user's question is embedded too, the closest chunks are retrieved, often reranked for relevance, and inserted into the prompt alongside the question. The LLM then generates its answer with that retrieved context in front of it.
Retrieval quality caps answer quality, an LLM can't ground an answer in a document that was never retrieved, so most RAG failures trace back to the retrieval step, not the model itself. Retrieval also adds real latency before generation even starts.
A RAG pipeline's total response time includes retrieval time on top of everything measured on a race card, the model's own inference time is only the last leg of a longer round trip.