Is the free, open-weight model actually a worse pick than the one you pay per token for? Averaged across all 25 tasks on this site: paid models (Haiku 4.5, Sonnet 5, Opus 4.8, Fable 5, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol) against every open-weight model on the site (Llama 4 Maverick, DeepSeek V3.2, Gemma 4 31B, GLM-5.2, Nemotron 3 Ultra, Nemotron 3.5 Lightning, MiniMax M3), recorded runs where they exist, clearly-labelled placeholder profiles where they don't, same as every other page here.
On this site's own numbers: paid models average faster, open-source models average cheaper (by roughly 449× per task), and paid models average more accurate. If that's not the outcome you expected, it isn't the one most marketing expects either. See the caveat below before treating it as a verdict.
Among all 14 models on this site, not just against each other:
The paid average here is built from 7 models spanning a genuine range (a fast/cheap tier through a slow/flagship-reasoning tier). The open-source average is built from 7 models across Meta, DeepSeek, Google, Z-AI, NVIDIA, MiniMax, a real spread of providers, not a single vendor's lineup. As more models get added and real runs replace placeholder profiles, this page recomputes automatically; it isn't hand-edited.
All figures are averages of this site's own per-task numbers: recorded monthly where a run exists, a clearly-labelled placeholder profile otherwise. See how we measure →. Want one specific pairing instead of an average? Compare AI Models runs any two models head-to-head on one task. For the full model-by-model breakdown, see the complete comparison →, or Claude vs ChatGPT → for the paid-vs-paid version of this question. New open-weight models land on the new models page → as soon as they're measured.