The intuition
A library, very quickly
Making Q, K and V
Same x, three lenses
Q = x WQ K = x WK V = x WV [T × d] · [d × dk] = [T × dk]
Every token gets all three
Matching and mixing
“it” looks for its antecedent
Mixing the values
outi = Σj wij · vj
Why keys and values are separate
Shapes cheat sheet
Self-attention vs cross-attention
Exceptions and sharp edges
Scores are not symmetric
read attention maps row by row: one row is one query.Tokens attend to themselves
nothing to fix, the residual connection relies on it too.dk must match
assert q.shape[-1] == k.shape[-1].Weights are not explanations
treat attention maps as hints, not proof.Bias or no bias
check the config before loading weights.