AI cost is what it actually costs to run an AI system in production, driven mainly by token usage (input and output, priced separately) multiplied by each model's per-token price, plus infrastructure and any added steps like retrieval or reranking.
Unlike most software, where compute cost is roughly fixed regardless of what a user asks, an LLM's cost scales directly with how much text goes in and comes out. A longer prompt costs more. A longer answer costs more. A more capable model usually costs more per token too.
The same task can cost a hundred times more on one model than another, ability, speed, and price don't move together, which is exactly why cost has to be measured per model per task, not assumed from a single average.
Providers charge separately, and at different rates, for input tokens and output tokens (output is typically pricier, since generating text is more computationally expensive than reading it). Total request cost is input tokens times input price, plus output tokens times output price, and a full pipeline's cost stacks every step that calls a model, retrieval, reranking, generation, each priced individually.
Cost, speed, and accuracy don't move together, the fastest model isn't always cheapest, and the cheapest isn't always least accurate, which is why real model selection needs all three numbers side by side, not one in isolation.
This is exactly the third number shown on every race card alongside speed and accuracy: real cost, computed from actual token counts and each model's published per-million-token price, not an estimate.