How Transformers Actually Work: Attention, Explained Intuitively

Abstract visualization of attention connections between tokens in a transformer neural network

Every large language model running today — from GPT to Llama to Claude — is built from the same core mechanism: self-attention. Introduced in 2017 by Vaswani and colleagues at Google in Attention Is All You Need, it solved a problem that had limited sequence models for decades. This article builds up the intuition behind queries, keys, and values, walks through the scaled dot-product formula with real numbers, explains why the square-root scaling exists, and shows what multiple attention heads actually learn.

The problem attention solved

Before 2017, sequence modelling was dominated by recurrent networks: LSTMs and GRUs processed one token at a time, carrying a hidden state forward. This had two costs. First, training was inherently sequential — you could not parallelize across a sequence the way GPUs demand. Second, information had to survive many sequential steps: in a 50-token sentence, connecting token 1 to token 50 meant a path through 50 recurrent transformations, and the signal decayed along the way.

Self-attention removes the path-length problem entirely. Every position in the sequence connects directly to every other position in a single operation, so long-range dependencies have the same cost as adjacent ones, and the whole sequence can be processed in parallel. The 2017 paper demonstrated the idea on machine translation, setting new state-of-the-art scores on the WMT 2014 English-to-German (28.4 BLEU) and English-to-French (41.8 BLEU) benchmarks — while training in days rather than weeks.

The search-engine intuition for Q, K, and V

Attention is easiest to understand as a soft, differentiable lookup:

  • Query (Q) — what this token is looking for. “Which other tokens are relevant to me?”
  • Key (K) — what this token offers as an index entry. “Here is what I contain.”
  • Value (V) — the actual content this token contributes if selected.

Think of a library. Your query is what you are searching for; each book has a key describing its contents on the spine; and the value is the book's content itself. You score every book's key against your query, pick the best matches with soft weights, and take a weighted blend of the corresponding contents.

In a transformer, these vectors are not hand-crafted. Each is a learned linear projection of the token's embedding, using weight matrices WQ, WK, and WV that are trained end to end. The model learns how to ask, how to index, and what to contribute.

Scaled dot-product attention, step by step

The whole operation is one formula: Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) · V

Here is what each step does:

  1. Score. Compute QK^T, the dot product of every query with every key. This gives an n × n score matrix, where entry (i, j) measures how similar query i is to key j. Larger means more relevant.
  2. Scale. Divide by sqrt(d_k), where dk is the dimension of the key vectors. The reason for this specific divisor is below — it is not arbitrary.
  3. Normalize. Apply the softmax to each row, turning raw scores into weights that sum to 1. Each row is now a probability distribution over the sequence.
  4. Blend. Multiply the weights by V. Each output position is a weighted average of all value vectors — one context-aware vector per token.

Let us walk through a tiny numerical example. Suppose a token has query q = [1.0, 0.5] and attends to two keys, k1 = [1.0, 0.2] and k2 = [0.2, 1.0], with values v1 = [1, 0] and v2 = [0, 1]. Here dk = 2, so we divide by sqrt(2) ≈ 1.414:

  • Scores: q·k1 = 1.10, q·k2 = 0.70
  • Scaled: 1.10/1.414 = 0.778, 0.70/1.414 = 0.495
  • Softmax: e^0.778 = 2.177, e^0.495 = 1.641; normalizing gives weights [0.570, 0.430]
  • Output: 0.570·[1,0] + 0.430·[0,1] = [0.570, 0.430]

The token's output is a blend: mostly the first value (it was more similar), but still informed by the second. In a real model, dk is typically 64 or 128 and n can be thousands, but the mechanics are exactly this.

Why divide by sqrt(dk)?

This is the one detail beginners most often gloss over, and it has a clean mathematical justification. Consider the idealized case where the components of q and k are independent random variables with mean 0 and variance 1. Then the variance of the dot product grows linearly with dimension — its typical magnitude grows as sqrt(dk).

Large scores are a problem because softmax is dominated by differences. A handful of slightly-larger logits grab nearly all the probability mass, and the rest get weights near zero — with near-zero gradients to match. Training stalls because almost nothing gets a useful update. A common misconception says scaling prevents softmax from going too flat; it is the opposite. Unscaled dot products make attention too peaked — one entry dominates. Dividing by sqrt(dk) cancels exactly the variance growth that dimensionality introduced, restoring unit variance so attention behaves the same way whether dk is 8 or 512.

Multi-head attention: specialists, not one generalist

Instead of computing attention once, the transformer runs it h times in parallel with separate learned projections:

MultiHead(Q, K, V) = Concat(head_1, ..., head_h) · W_O

head_i = Attention(Q·W_i^Q, K·W_i^K, V·W_i^V)

Each head projects Q, K, and V into a smaller subspace (dk = dmodel/h), performs attention there, and the results are concatenated and mixed by the output projection WO. The original paper used 8 heads with dmodel = 512, so each head worked in 64 dimensions.

The intuition: one attention mapping has to capture everything — syntax, semantics, position, coreference — in a single set of projections. Multiple heads can specialize. One head might track subject–verb links, another pronoun references, another positional adjacency. And the cost stays flat: the per-head dimension shrinks as heads are added, so 8 heads at dk = 64 use roughly the compute of one head at dmodel = 512.

This specialization is not just a hypothesis. Clark and colleagues (2019) showed BERT heads linking verbs to direct objects and resolving coreference; Vig and Belinkov (2019) found GPT-2 heads tracking list items, acronyms, and paired punctuation; Elhage and colleagues (2021) described “induction heads” that implement pattern completion — seeing [A][B]…[A] and boosting the probability of [B]. One fair caveat: attention weights are clues, not proof of what the model decided. They are worth inspecting, but causal experiments — not heatmaps — are what establish mechanism.

Where attention sits in the architecture

Attention is not the whole transformer; it is one ingredient, used in three distinct roles:

  • Encoder self-attention. Every token attends to every other token, bidirectionally. This is what BERT later exploited for masked language modelling.
  • Decoder masked self-attention. A causal mask (adding −∞ to future positions before the softmax) prevents the model from peeking at tokens it has not generated yet. Modern LLMs like GPT and Llama are decoder-only: this masked attention is the entire architecture.
  • Cross-attention. The decoder's queries attend to the encoder's keys and values, letting generated tokens draw on the input sequence.

Two more pieces make it work. Self-attention is order-blind — shuffle the input and nothing changes — so the original paper added sinusoidal positional encodings to the input embeddings, giving each position a unique, distance-aware signature. (Modern models mostly use learned or rotary embeddings instead, but the need is the same.) Each attention block is also wrapped in residual connections and layer normalization: LayerNorm(x + Sublayer(x)), which keeps gradients flowing through the stack.

The price: quadratic cost

Attention's power has a cost. The n × n score matrix grows quadratically with sequence length, in both compute and memory. A 2,000-token context (GPT-3's original window) means 4 million scores per head per layer; a 128,000-token context means 16 billion. This quadratic bottleneck is why long-context inference is hard and why so much modern research targets it.

The most influential efficiency advance is FlashAttention (Dao et al., 2022): an exact — not approximate — reordering of the computation that tiles the score matrix so it never needs to be fully materialized, cutting memory from O(n²) to O(n).

Putting it together

A transformer block, then, is a rhythm repeated dozens of times: normalize, gather context with multi-head attention, add back the residual, normalize again, reason per position with a feed-forward network, add back the residual. The 2017 base model used 6 such blocks on the encoder side and 6 on the decoder, with dmodel = 512, 8 heads, a feed-forward width of 2048, dropout of 0.1, and about 65 million parameters — trained with Adam and a 4,000-step learning-rate warmup. Everything since has been variations on this theme: more layers, bigger dimensions, decoder-only, better training — but the attention core is unchanged.

The deep idea is simple enough to state in one sentence: let every token dynamically decide, for itself, which other tokens matter — and build its representation from exactly that. Queries, keys, and values are just the machinery that makes the deciding differentiable.

Further reading

Similar Posts

Leave a Reply