Inference is the act of running a trained model to produce an output, generating a response to a prompt, as opposed to training, which is the earlier, much more expensive process of teaching the model in the first place.
Training happens once (or occasionally, when a new model version is built). Inference happens every single time someone sends a prompt and gets a response back. It's the part that actually runs in production, and the part every speed and cost number on this site is measuring.
Training cost is a one-time (if very large) expense paid by the model provider. Inference cost is paid on every single request, forever, which is why inference speed and cost, not training cost, is what shapes the economics of an AI product.
A prompt is tokenized, fed through the model's layers to compute the next token's probabilities (prefill), and that process repeats one token at a time until the response is complete (decode). The model's weights themselves never change during this, only the input and the internal state built up while processing it.
Inference speed and cost scale with output length, model size, and hardware, the three levers every AI provider and every AI product team is actually optimizing.
This is exactly what every race on this site times: real inference, start to finish, on the provider's own infrastructure, not an estimate.