Decode

Inference & Serving Advanced

Decode is the phase of inference that generates the response one token at a time, each new token depending on every token that came before it, including the ones just generated.

In simple terms

After prefill finishes reading the prompt, decode starts writing the answer, but strictly one word-piece at a time: generate a token, feed it back in, generate the next one, repeat. That's why a longer answer takes proportionally longer, unlike prefill which processes the whole prompt at once.

Why it matters

Decode can't be parallelized across tokens the way prefill can, each token needs the previous one to exist first, so it's the phase where the KV cache, batching strategy, and raw memory bandwidth of the serving hardware matter most.

How it works

For each step, the model computes attention for exactly one new token against the entire cached history, produces a probability distribution over the vocabulary, and samples (or picks the most likely) next token. That token is appended to the sequence and the cache, and the process repeats until an end-of-sequence token or length limit is reached.

Where it fits

KV cache (from prefill) → Decode step: generate one token → Append to cache → Decode step (repeat) → End-of-sequence token → Complete response

Production impact

Decode is memory-bandwidth-bound, not compute-bound, so tokens-per-second is largely capped by how fast the GPU can read the growing KV cache, which is why techniques like continuous batching (serving many requests' decode steps together) exist to keep GPUs busy.

On howfastai.com

Once the first token appears, the pace of everything after it, this site's generation-speed number, is decode speed.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →