Decode is the phase of inference that generates the response one token at a time, each new token depending on every token that came before it, including the ones just generated.
After prefill finishes reading the prompt, decode starts writing the answer, but strictly one word-piece at a time: generate a token, feed it back in, generate the next one, repeat. That's why a longer answer takes proportionally longer, unlike prefill which processes the whole prompt at once.
Decode can't be parallelized across tokens the way prefill can, each token needs the previous one to exist first, so it's the phase where the KV cache, batching strategy, and raw memory bandwidth of the serving hardware matter most.
For each step, the model computes attention for exactly one new token against the entire cached history, produces a probability distribution over the vocabulary, and samples (or picks the most likely) next token. That token is appended to the sequence and the cache, and the process repeats until an end-of-sequence token or length limit is reached.
Decode is memory-bandwidth-bound, not compute-bound, so tokens-per-second is largely capped by how fast the GPU can read the growing KV cache, which is why techniques like continuous batching (serving many requests' decode steps together) exist to keep GPUs busy.
Once the first token appears, the pace of everything after it, this site's generation-speed number, is decode speed.