Inference

Inference & Serving Beginner

Inference is the act of running a trained model to produce an output, generating a response to a prompt, as opposed to training, which is the earlier, much more expensive process of teaching the model in the first place.

In simple terms

Training happens once (or occasionally, when a new model version is built). Inference happens every single time someone sends a prompt and gets a response back. It's the part that actually runs in production, and the part every speed and cost number on this site is measuring.

Why it matters

Training cost is a one-time (if very large) expense paid by the model provider. Inference cost is paid on every single request, forever, which is why inference speed and cost, not training cost, is what shapes the economics of an AI product.

How it works

A prompt is tokenized, fed through the model's layers to compute the next token's probabilities (prefill), and that process repeats one token at a time until the response is complete (decode). The model's weights themselves never change during this, only the input and the internal state built up while processing it.

Where it fits

Prompt → Tokenization → Prefill → Decode (token by token) → Detokenization → Response

Production impact

Inference speed and cost scale with output length, model size, and hardware, the three levers every AI provider and every AI product team is actually optimizing.

On howfastai.com

This is exactly what every race on this site times: real inference, start to finish, on the provider's own infrastructure, not an estimate.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →