An LLM is a neural network trained on huge amounts of text to predict the next word in a sequence, well enough that stringing those predictions together produces coherent answers, summaries, and code.
Think of it as a very large, very well-read autocomplete. Given the text so far, it estimates the probability of every possible next word and picks one. Do that one word at a time, thousands of times a second, and the output reads like a written answer.
Every product this site measures, chat assistants, coding tools, voice agents, is an LLM (or several) wrapped in application logic. Understanding what the model actually does underneath is the difference between debugging a real system and treating it as a black box.
An LLM is a transformer trained in two stages: pretraining, where it learns language and world knowledge from raw text by predicting missing or next words, and fine-tuning, where it's adjusted to follow instructions and hold a conversation. At inference time, the trained weights never change, only the text going in and coming out.
Model size and architecture set the ceiling on speed, cost, and accuracy at the same time, a bigger model is usually slower and pricier per token, but often more accurate, which is exactly the trade-off this site measures per task.
Every race on this site is one LLM answering one task. The multiplier, cost, and accuracy numbers shown are that model's actual measured behavior, not a spec sheet claim.