Prefill is the first phase of inference, processing the entire input prompt in parallel to build up the KV cache, before the model generates a single token of the response.
Before an LLM writes anything, it has to "read" your whole prompt first. That reading step, computing attention over every prompt token at once, is prefill. It finishes right before the first word of the answer appears.
Prefill time is most of what "time to first token" measures. A long prompt takes longer to prefill, which is why a request with a huge attached document takes noticeably longer to start responding than a short question.
The entire tokenized prompt is fed through the model's layers in one parallel pass, computing and caching the key and value vectors for every prompt token. Because this can happen in parallel across the whole prompt (unlike decode, which is strictly sequential), prefill is compute-bound and benefits heavily from raw GPU throughput.
Prefill is compute-bound while decode is memory-bandwidth-bound, meaning the two phases stress a GPU differently, which is exactly why serving systems batch and schedule them separately for efficiency.
Time to first token, one of the core numbers this site measures per task, is essentially prefill time plus queueing, before any generation speed comes into play at all.