Transformer

LLM Foundations Intermediate

The transformer is the neural network architecture nearly every modern LLM is built on, defined by its use of self-attention to weigh every token against every other token at once, instead of processing text strictly left to right.

In simple terms

Older architectures read a sentence one word at a time and tried to remember what came before. A transformer instead looks at the whole sequence simultaneously and directly compares each word against every other word to figure out which ones matter to each other. That comparison step is attention.

Why it matters

This one design choice is why LLMs scaled the way they did: attention is easy to parallelize across GPU cores, which made it practical to train on far more data and far bigger models than earlier architectures allowed.

How it works

A transformer stacks many identical layers, each combining a self-attention block (tokens comparing themselves to each other) with a small feed-forward network (per-token processing), plus normalization and residual connections that keep training stable at depth. Most LLMs use only the decoder half of the original transformer design, built to generate text one token at a time.

Where it fits

Token embeddings → Positional encoding → Transformer layers (attention + MLP) → Output probabilities → Next token

Production impact

Transformer depth and width (how many layers, how wide each layer is) is most of what "model size" means, and size trades directly against inference speed: more layers means more GPU work per token generated.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →