Paid AI vs open-source AI

Is the free, open-weight model actually a worse pick than the one you pay per token for? Averaged across all 25 tasks on this site: paid models (Haiku 4.5, Sonnet 5, Opus 4.8, Fable 5, GPT-5.6 Luna, GPT-5.6 Terra, GPT-5.6 Sol) against every open-weight model on the site (Llama 4 Maverick, DeepSeek V3.2, Gemma 4 31B, GLM-5.2, Nemotron 3 Ultra, Nemotron 3.5 Lightning, MiniMax M3), recorded runs where they exist, clearly-labelled placeholder profiles where they don't, same as every other page here.

The averages, side by side

Paid (7 models)
Open-source (7 models)
Speed
444× faster than a person
395× faster than a person
Cost
$0.03 per task
$0.0001 per task
Accuracy
88%
83%

On this site's own numbers: paid models average faster, open-source models average cheaper (by roughly 449× per task), and paid models average more accurate. If that's not the outcome you expected, it isn't the one most marketing expects either. See the caveat below before treating it as a verdict.

Where the open-source models actually rank

Among all 14 models on this site, not just against each other:

  • Llama 4 Maverick (Meta): #6 fastest, #7 cheapest, 81% average accuracy, 484× faster than a person on average
  • DeepSeek V3.2 (DeepSeek): #11 fastest, #6 cheapest, 89% average accuracy, 257× faster than a person on average
  • Gemma 4 31B (Google): #3 fastest, #1 cheapest, 77% average accuracy, 678× faster than a person on average
  • GLM-5.2 (Z-AI): #9 fastest, #2 cheapest, 88% average accuracy, 306× faster than a person on average
  • Nemotron 3 Ultra (NVIDIA): #13 fastest, #3 cheapest, 87% average accuracy, 162× faster than a person on average
  • Nemotron 3.5 Lightning (NVIDIA): #5 fastest, #4 cheapest, 76% average accuracy, 489× faster than a person on average
  • MiniMax M3 (MiniMax): #8 fastest, #5 cheapest, 84% average accuracy, 386× faster than a person on average

Best pick per difficulty tier, every provider

Anthropic
OpenAI
Meta
DeepSeek
Google
Z-AI
NVIDIA
MiniMax
Simple
Haiku 4.5
GPT-5.6 Luna
Llama 4 Maverick
Llama 4 Maverick
Gemma 4 31B
Gemma 4 31B
Nemotron 3.5 Lightning
Gemma 4 31B
Intermediate
Sonnet 5
GPT-5.6 Terra
Llama 4 Maverick
DeepSeek V3.2
Gemma 4 31B
GLM-5.2
Nemotron 3 Ultra
MiniMax M3
Complex
Opus 4.8
GPT-5.6 Sol
DeepSeek V3.2
DeepSeek V3.2
GLM-5.2
GLM-5.2
Nemotron 3 Ultra
MiniMax M3

The honest caveat

The paid average here is built from 7 models spanning a genuine range (a fast/cheap tier through a slow/flagship-reasoning tier). The open-source average is built from 7 models across Meta, DeepSeek, Google, Z-AI, NVIDIA, MiniMax, a real spread of providers, not a single vendor's lineup. As more models get added and real runs replace placeholder profiles, this page recomputes automatically; it isn't hand-edited.

All figures are averages of this site's own per-task numbers: recorded monthly where a run exists, a clearly-labelled placeholder profile otherwise. See how we measure →. Want one specific pairing instead of an average? Compare AI Models runs any two models head-to-head on one task. For the full model-by-model breakdown, see the complete comparison →, or Claude vs ChatGPT → for the paid-vs-paid version of this question. New open-weight models land on the new models page → as soon as they're measured.