The transformer is the neural network architecture nearly every modern LLM is built on, defined by its use of self-attention to weigh every token against every other token at once, instead of processing text strictly left to right.
Older architectures read a sentence one word at a time and tried to remember what came before. A transformer instead looks at the whole sequence simultaneously and directly compares each word against every other word to figure out which ones matter to each other. That comparison step is attention.
This one design choice is why LLMs scaled the way they did: attention is easy to parallelize across GPU cores, which made it practical to train on far more data and far bigger models than earlier architectures allowed.
A transformer stacks many identical layers, each combining a self-attention block (tokens comparing themselves to each other) with a small feed-forward network (per-token processing), plus normalization and residual connections that keep training stable at depth. Most LLMs use only the decoder half of the original transformer design, built to generate text one token at a time.
Transformer depth and width (how many layers, how wide each layer is) is most of what "model size" means, and size trades directly against inference speed: more layers means more GPU work per token generated.