Plain definitions for the terms used across this site — each one short enough to quote on its own, and checked against how this site actually measures things, not a generic dictionary gloss.
Time to first token is how long an AI model takes to start responding after receiving a prompt: the gap between sending the request and the first word appearing, measured in milliseconds. For models that show their reasoning as they work, thinking tokens count the same as answer tokens in that measurement, unless a model's own timing data separates the two.
A token is the unit AI providers use to measure and bill text: roughly three-quarters of a word in English, though it varies by language and model. Every prompt sent is counted as input tokens; every reply the AI writes is counted as output tokens, and providers usually charge a different rate for each.
Tokens per second measures how quickly an AI model writes its answer once it starts, not how long it took to start (that is time to first token). A model can have a fast first-token time but slow generation speed, or the reverse: the two numbers describe different parts of the same response.
Accuracy cannot be read from an API response, so it has to be judged: each AI answer is checked against a fixed grading rubric written for that specific task, either by a human reviewer or an AI judge model, and which method produced the score is always disclosed alongside it.
Speed × Quality is a blended ranking score, not a literal multiplication despite the name: thirty percent normalized speed and seventy percent normalized accuracy, combined into one number. It exists because the fastest AI model and the most accurate one are rarely the same model; the score picks a winner by weighting accuracy over raw speed.
Extended thinking is a mode some AI models use to reason step by step before answering, adding a visible delay before the final response starts streaming. A small subset of recorded runs used it with a separate thinking-time measurement; it is not the standard recording mode for every model on this site.
A recorded run is a real request sent to an AI provider on a fixed monthly schedule, with the response streamed back and every word timestamped as it arrived. Nothing is generated live while a visitor looks at a result: pressing Race or Run replays that recording at the pace it originally happened.
A placeholder profile is a clearly labeled stand-in used where a real recorded run or accuracy score does not exist yet, so a page still works end to end. It is never blended with real data or presented as measured: every result states outright whether it is a recorded number or a placeholder.
A model tier, Simple, Intermediate or Complex, describes how demanding a task is, not how capable a specific AI model is. A Simple tag does not mean every model handles it easily, and a Complex tag does not mean every model struggles: the tier is paired against a model's own strengths to produce the fit verdict shown on each result.
← Discover AI Speed · How we measure → · FAQ → · Voice AI → · Is 10 tok/s fast? →