An embedding is a list of numbers, typically hundreds or thousands of them, that represents the meaning of a piece of text (or image, or audio) as a point in space, so that similar meanings end up as nearby points.
Imagine plotting every sentence you've ever read as a dot on a map, where sentences about similar things land close together. An embedding model does that automatically: it reads text and outputs coordinates. "Dog" and "puppy" land near each other; "dog" and "spreadsheet" land far apart.
Embeddings are what make semantic search possible, finding text by meaning instead of exact keyword match. Without them, a search for "how do I cancel my subscription" would miss a help article titled "ending your plan."
A smaller, specialized neural network (often itself a transformer) reads the input and produces a fixed-length vector. That vector is trained so that texts with similar meaning land close together by some distance measure, usually cosine similarity, and dissimilar texts land far apart.
Embedding a document is a one-time (or occasional-update) cost; embedding a user's query happens on every search, so its latency sits directly in the critical path of any RAG system.