Transformers, from tokens to attention
Start with a short sequence. Turn it into vectors. Then let each position gather information from the others. That is the operation at the heart of a transformer.
On this page
Tokens need context
Consider the phrase “the river bank.” A token representation becomes more useful when it can incorporate nearby context: “river” helps distinguish this use of “bank” from a financial institution. For a first walkthrough, imagine each word is one token, although real tokenizers often split words into smaller pieces.
An embedding maps each token ID to a vector. Position information gives the model a way to distinguish order. Self-attention then mixes information across positions; the output still has one vector per input position.
Queries, keys, and values
Three learned projections of the input produce queries, keys, and values. A query is compared with every allowed key. The resulting scores determine how much of each value contributes to that position’s output.
- Query: the vector used to ask for relevant context.
- Key: the vector compared against a query.
- Value: the information mixed into the result.
Dividing by the square root of the key dimension controls the scale of the scores. For one sequence with T positions, the score matrix is T × T: every row describes the context gathered for one query position.
A small causal-attention example
In next-token prediction, a position must not see future tokens. A causal mask excludes those keys before softmax. During training, this lets us compute many positions together without revealing their future context.
import torch
import torch.nn.functional as F
torch.manual_seed(7)
# Batch, heads, sequence length, head dimension
q = torch.randn(1, 1, 3, 8)
k = torch.randn(1, 1, 3, 8)
v = torch.randn(1, 1, 3, 8)
context = F.scaled_dot_product_attention(
q, k, v,
is_causal=True,
dropout_p=0.0,
)
print(context.shape) # torch.Size([1, 1, 3, 8])Here q, k, and v are random tensors so the example can run on its own. In a model, learned projections create them from hidden states. The function performs scaling, masking, softmax, and value aggregation; it does not supply those projections or a complete transformer block.
Beyond one attention head
Multi-head attention runs several learned projections in parallel and combines their outputs. A transformer layer also uses a feed-forward network, residual connections, and normalization. Encoder-decoder models add a route for the decoder to attend to the encoder’s outputs.
When inspecting an implementation, follow the shapes before tuning anything. Write down the batch size, sequence length, model width, and number of heads. Then check which dimension softmax uses and which positions the mask allows. This gives you a concrete way to compare the diagram, the equation, and the code.
Sources & further reading
Primary references for the concepts and APIs in this article.