Reranking is a second, more precise scoring pass over a retrieval system's initial results, reordering them by relevance before the top few are handed to the LLM, catching cases the first-pass search got wrong.
Initial retrieval is built for speed: it has to scan a huge collection fast, so it uses a cheaper relevance signal. A reranker is slower but smarter, it looks closely at just the candidates retrieval already narrowed down (maybe the top 50) and re-scores them properly, then keeps only the best handful.
Fast retrieval and precise relevance scoring are in tension, doing the precise version over an entire document collection would be too slow. Reranking gets both: broad, fast recall first, then narrow, careful precision second.
A cross-encoder model reads the query and each candidate document together (not separately, the way embeddings do) and outputs a direct relevance score. Because it processes query and document jointly, it's far more accurate than comparing pre-computed embeddings, but also too slow to run over an entire collection, so it only ever sees retrieval's already-narrowed shortlist.
Reranking adds a real latency cost, one model call per candidate pair, but usually improves answer groundedness enough to be worth it in any RAG system where accuracy matters more than shaving off the last few hundred milliseconds.