Vector Database

Embeddings & RAG Intermediate

A vector database stores embeddings and answers "which of these millions of vectors are closest to this one" quickly, using specialized indexing rather than comparing against every stored vector one by one.

In simple terms

A regular database is good at finding rows that exactly match a value. A vector database is built for a different question: given a point in space, which stored points are nearest to it. That's a much harder problem to make fast at scale, which is why it gets its own kind of database.

Why it matters

Comparing a query embedding against every document embedding one at a time (brute force) doesn't scale past a small collection. Vector databases use approximate nearest-neighbor indexes to answer the same question in milliseconds across millions of vectors, trading a small amount of accuracy for a large amount of speed.

How it works

Most vector databases build a graph or tree-like index over the stored vectors (HNSW is the most common approach today) so a search only has to check a small fraction of the collection instead of all of it. Many also support filtering by metadata (date, category, source) alongside the similarity search, and sharding or replication to scale beyond one machine.

Where it fits

Document embeddings → Vector database index → Query embedding → Approximate nearest-neighbor search → Top-K results

Production impact

Index choice trades recall against speed and memory: a more thorough index finds better matches but costs more RAM and more time per query, a real tuning decision in any production RAG system.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →