AI terminology

Search by name, acronym, or category. Filter by how deep it goes. Every term links to what to learn before it and after it.

LLM Foundations

What a language model actually is, and the mechanics underneath it.

Embeddings & RAG

Giving a model access to information it was never trained on.

Agents & Tools

Models that take actions, not just answer questions.

Inference & Serving

What happens between a prompt and a response, and how it's made fast.

Decode

Decode is the phase of inference that generates the response one token at a time, each new token depending on every token that came before it, including the ones just generated.

Advanced

Inference

Inference is the act of running a trained model to produce an output, generating a response to a prompt, as opposed to training, which is the earlier, much more expensive process of teaching the model in the first place.

Beginner

KV Cache

The KV cache stores the key and value vectors a transformer computes for every token it has already processed, so generating the next token can reuse that work instead of recomputing attention over the whole sequence from scratch.

Advanced

Model Serving

Model serving is the infrastructure layer that takes a trained model and makes it available to answer real requests, handling batching, queueing, and GPU allocation so many requests can be served efficiently at once.

Intermediate

Prefill

Prefill is the first phase of inference, processing the entire input prompt in parallel to build up the KV cache, before the model generates a single token of the response.

Advanced

Quantization

Quantization reduces the numeric precision used to store a model's weights, for example from 16-bit to 8-bit or 4-bit numbers, shrinking memory use and often speeding up inference, at some cost to accuracy.

Advanced

vLLM

vLLM is an open-source inference engine built around paged attention, a technique that manages the KV cache in fixed-size blocks (like an operating system manages memory pages) to dramatically reduce wasted GPU memory.

Infrastructure

AI Infrastructure

The hardware and routing layer underneath every request.

AI Operations

Running AI systems in production: watching them, measuring them, keeping them up.

AI Security & Cost

What can go wrong, and what it costs.

← Knowledge map · Discover AI Speed →