Attention

Self-Attention LLM Foundations Advanced

Attention is the mechanism a transformer uses to decide, for each token, how much every other token in the sequence should influence it, letting the model connect a pronoun to the noun it refers to, or a closing bracket to the one it opens.

In simple terms

For every token, the model asks three questions of every other token: what am I looking for (query), what do I offer (key), and what information do I actually carry (value). It scores how well each query matches each key, turns those scores into weights, and blends the values accordingly. Do that for every token against every other token, and the model ends up with a rich, context-aware representation of each word.

Why it matters

Attention is what lets a model handle long-range dependencies, a word on page one affecting a word on page ten, without the information degrading the way it did in older sequential architectures. It's also the single most expensive part of inference to compute and store.

How it works

Each token's embedding is projected into a query, key, and value vector. Attention scores are the dot product of a token's query against every other token's key, scaled and passed through a softmax to become weights that sum to one. Those weights combine the value vectors into the token's new representation. Multi-head attention runs several of these in parallel with different learned projections, so the model can track several kinds of relationships at once.

Where it fits

Token embedding → Query, Key, Value projections → Attention scores → Weighted sum of values → Updated token representation

Production impact

Attention's compute and memory cost grows with the square of sequence length in the naive form, which is exactly why long-context requests are slower and pricier, and why techniques like KV caching and Flash Attention exist.

Learn this first

Related terms

Learn next

← All terms · Knowledge map →