Every number on this site is either recorded, scored, or a clearly labelled placeholder. Nothing is measured live while you use the site.
Recorded. For each of the 25 tasks we send the same prompt to each of the 14 models (across Anthropic, OpenAI, Meta, DeepSeek, Google, Z-AI, NVIDIA, and MiniMax) and record three real timestamps off the request itself: time to the first token (the model reading and understanding the request), generation time (writing or working out the answer, measured from that first token to the last), and total wall-clock time. Token counts come from each provider's own API usage field, never estimated from a word or character count. Runs recorded so far: 193.
What "time to first token" actually captures. It's the moment the first piece of content streams back — for a model that shows its reasoning as it happens, a reasoning token counts the same as an answer token, so most runs don't separate "thinking" from "answering" inside that number. A small subset of Anthropic runs were recorded with extended thinking explicitly enabled and do carry a separate thinking-time field on top of the usual three; that's not our standard recording mode, so treat it as present-when-shown rather than a site-wide guarantee.
Scored. Accuracy cannot be measured from an API. Each task breaks into stages (reading the request, understanding it, writing the answer, and so on), each with its own weight (the weights add to 100) and its own accuracy score, either a real human-reviewed or AI-judged score when one exists, or a labelled placeholder profile when it doesn't, always stated which. The accuracy you see is the weighted average across a task's stages. Speed × Quality is a 30/70 blend of normalised speed and quality, not a literal product, despite the label. Scores so far: 627. AI-judged scores are graded by Opus 4.8, one of the models this site measures, so it is sometimes grading its own answers. That's a real conflict of interest, not hidden: a human-reviewed calibration set is how we check the judge isn't just favouring itself, and it's still being built. Until it's complete, treat AI-judged numbers as a real measurement, not yet an independently verified one.
Tokens & cost. Every result also shows tokens used and an estimated dollar cost. The cost is tokens × each model's current published price per million tokens (input and output priced separately, since they differ). Token counts come from a recorded run when one exists; otherwise the same placeholder token count is used for every model on that task, so the dollar figure you see is driven by price differences between models, not by noise in the token estimate. That's usually the biggest surprise on the page: the same task, the same rough token count, a wildly different bill.
Why "X times faster" claims need a second look. In September 2026, a startup called TypeSafe AI released a model called Jev built to return a fixed-shape decision (a classification plus a confidence score) instead of writing text, and reported it running 20 to 200 times faster than models like Claude or GPT. That's a real number, but it's not one this site can put next to Sonnet or Opus on a chart: every model we measure writes an open-ended reply to a prompt, and Jev structurally can't. Comparing the two is like timing a light switch against a light bulb: both do something fast, just not the same something. Whenever you see a speed claim anywhere, including here, the useful question isn't "how fast" on its own, it's "fast at doing exactly what."
Tiers. Every task is tagged Simple, Intermediate, or Complex: roughly everyday, harder, and specialist work. The tag describes the task, not the model: a Simple tag doesn't mean every model breezes through it, and a Complex tag doesn't mean every model struggles. Pairing a task's tier against a model's own tier is what the green/yellow/red "fit" verdict on each result is judging.
Placeholders. Where a run or score does not exist yet, the site shows a labelled placeholder profile so the experience works end to end. Placeholders are replaced, never blended.
"You". The human time is a typical estimate for a person doing the task by hand. It is an estimate, and it says so on the card. Take the 15-second typing test under the "You" gauge and it becomes real: your own words-per-minute replaces the writing portion of that estimate (the reading and thinking portions stay generic — a 15-second typing test measures typing, not comprehension), clamped to a believable range so a fluke sprint or a distracted near-empty attempt can't swing a task's time by 10×. Nothing typed leaves your browser; the result is stored locally and can be reset from the same panel.
What we deliberately don't do. Each task × model pair is recorded once, not averaged or median'd across repeated attempts — a single real sample, not a smoothed one. There's no throwaway warm-up request run first, and no automatic retry or outlier check beyond whatever a provider's own SDK does by default; if a request fails outright, it's skipped and picked up on the next monthly cycle, not silently substituted. There's also no technical check that stops a host from serving a request unusually fast or slow — the real safeguard here is disclosure, not detection: every run records which host actually served it (81 via openrouter, 48 via openai, 45 via anthropic — 19 older runs predate this field and don't say), so a number can't quietly look faster than it really was without saying where it came from.
← Discover AI Speed · Compare AI Models → · FAQ → · Model comparison → · New models → · Voice AI → · Glossary → · Open dataset → · Is 10 tok/s fast? →