GPU

Graphics Processing Unit AI Infrastructure Beginner

A GPU is a processor built with thousands of small cores designed to do the same simple operation on many pieces of data at once, which happens to be exactly the kind of math (matrix multiplication) that neural networks are built from.

In simple terms

A CPU is a small number of powerful cores good at doing many different things in sequence. A GPU is the opposite: thousands of simpler cores good at doing the exact same operation to huge amounts of data simultaneously. Training and running an LLM is almost entirely that kind of repetitive, parallel math.

Why it matters

GPUs (originally built for rendering graphics) turned out to be close to perfectly suited for neural network math, which is why they, not CPUs, became the default hardware for both training and serving every modern LLM.

How it works

A GPU's cores execute the same instruction across many data elements in lockstep (single instruction, multiple data), and its high-bandwidth memory feeds those cores fast enough to keep them busy. Model weights and activations live in GPU memory (VRAM/HBM), and matrix multiplications, the core operation in a transformer, are dispatched across thousands of cores at once.

Where it fits

Model weights (loaded into GPU memory) → Incoming request → GPU cores: parallel matrix multiplication → Output tokens

Production impact

GPU memory capacity limits how large a model (or how many concurrent requests) a single GPU can serve, and memory bandwidth, not just raw compute, is often the actual bottleneck during decode.

On howfastai.com

Every model measured on this site runs on GPUs somewhere behind the provider's API, and which GPU, how many, and how well the serving stack uses them shapes the seconds shown on every race card just as much as the model's own size does.

Related terms

Learn next

← All terms · Knowledge map →