The KV cache stores the key and value vectors a transformer computes for every token it has already processed, so generating the next token can reuse that work instead of recomputing attention over the whole sequence from scratch.
Without caching, generating each new token would mean re-running attention over every previous token, from the start, every single time, wastefully redoing work that hasn't changed. The KV cache remembers each token's key and value vectors the first time they're computed, so decode only has to do the new token's share of the work.
It's the single biggest reason modern LLM inference is as fast as it is; without it, generating a long response would get progressively, drastically slower as the sequence grows, instead of staying roughly constant per token.
During prefill, the model computes key and value vectors for every prompt token and stores them in GPU memory. During decode, each new token only needs to compute its own query, key, and value, then attends against the entire cached history instead of recomputing it, and its own key and value get appended to the cache for the next step.
The KV cache grows with sequence length and consumes GPU memory directly, on long-context requests it can be the actual limit on how many concurrent requests a GPU can serve, more than compute is. This is why techniques like paged attention and cache quantization exist.
Once a response starts streaming, how fast it continues (tokens per second) is largely a function of how efficiently the serving stack manages this cache, not just raw GPU speed.