← All writing
Deep learningSeptember 20263 min read

Transformers, from tokens to attention

Start with a short sequence. Turn it into vectors. Then let each position gather information from the others. That is the operation at the heart of a transformer.

On this page

Tokens need context

Consider the phrase “the river bank.” A token representation becomes more useful when it can incorporate nearby context: “river” helps distinguish this use of “bank” from a financial institution. For a first walkthrough, imagine each word is one token, although real tokenizers often split words into smaller pieces.

An embedding maps each token ID to a vector. Position information gives the model a way to distinguish order. Self-attention then mixes information across positions; the output still has one vector per input position.

A simplified decoder-only language model. Each block includes residual connections and normalization, omitted here to keep the main path readable.

Queries, keys, and values

Three learned projections of the input produce queries, keys, and values. A query is compared with every allowed key. The resulting scores determine how much of each value contributes to that position’s output.

  • Query: the vector used to ask for relevant context.
  • Key: the vector compared against a query.
  • Value: the information mixed into the result.
Attention(Q, K, V) = softmax(QKᵀ / √dₖ)V
dₖ is the key dimension. Softmax runs across the key positions in each row; the resulting weights multiply V.

Dividing by the square root of the key dimension controls the scale of the scores. For one sequence with T positions, the score matrix is T × T: every row describes the context gathered for one query position.

An illustrative attention row for “bank”: 0.10 × value(the) + 0.65 × value(river) + 0.25 × value(bank). These weights are made up, not measurements from a trained model.

A small causal-attention example

In next-token prediction, a position must not see future tokens. A causal mask excludes those keys before softmax. During training, this lets us compute many positions together without revealing their future context.

One attention head · PyTorchpython
import torch
import torch.nn.functional as F

torch.manual_seed(7)
# Batch, heads, sequence length, head dimension
q = torch.randn(1, 1, 3, 8)
k = torch.randn(1, 1, 3, 8)
v = torch.randn(1, 1, 3, 8)

context = F.scaled_dot_product_attention(
    q, k, v,
    is_causal=True,
    dropout_p=0.0,
)
print(context.shape)  # torch.Size([1, 1, 3, 8])

Here q, k, and v are random tensors so the example can run on its own. In a model, learned projections create them from hidden states. The function performs scaling, masking, softmax, and value aggregation; it does not supply those projections or a complete transformer block.

Beyond one attention head

Multi-head attention runs several learned projections in parallel and combines their outputs. A transformer layer also uses a feed-forward network, residual connections, and normalization. Encoder-decoder models add a route for the decoder to attend to the encoder’s outputs.

When inspecting an implementation, follow the shapes before tuning anything. Write down the batch size, sequence length, model width, and number of heads. Then check which dimension softmax uses and which positions the mask allows. This gives you a concrete way to compare the diagram, the equation, and the code.

Sources & further reading

Primary references for the concepts and APIs in this article.

  1. Vaswani et al. — Attention Is All You Need (opens in a new tab)
  2. PyTorch — Scaled dot product attention (opens in a new tab)
  3. PyTorch — Transformer reference (opens in a new tab)