A GPU is a processor built with thousands of small cores designed to do the same simple operation on many pieces of data at once, which happens to be exactly the kind of math (matrix multiplication) that neural networks are built from.
A CPU is a small number of powerful cores good at doing many different things in sequence. A GPU is the opposite: thousands of simpler cores good at doing the exact same operation to huge amounts of data simultaneously. Training and running an LLM is almost entirely that kind of repetitive, parallel math.
GPUs (originally built for rendering graphics) turned out to be close to perfectly suited for neural network math, which is why they, not CPUs, became the default hardware for both training and serving every modern LLM.
A GPU's cores execute the same instruction across many data elements in lockstep (single instruction, multiple data), and its high-bandwidth memory feeds those cores fast enough to keep them busy. Model weights and activations live in GPU memory (VRAM/HBM), and matrix multiplications, the core operation in a transformer, are dispatched across thousands of cores at once.
GPU memory capacity limits how large a model (or how many concurrent requests) a single GPU can serve, and memory bandwidth, not just raw compute, is often the actual bottleneck during decode.
Every model measured on this site runs on GPUs somewhere behind the provider's API, and which GPU, how many, and how well the serving stack uses them shapes the seconds shown on every race card just as much as the model's own size does.