The lookup
Seven ids fetch seven rows
Why it counts as a matrix multiply
A map of meaning
Neighbours in embedding space
Vectors as colour strips
Cosine similarity between them
Cosine similarity is an angle
cos θ = a · b ÷ (‖a‖ ‖b‖)
Where the map comes from
Training pulls the points into clusters
The table is big, so some models share it
Exceptions and sharp edges
One vector per token, any context
attention mixes in context later. That is the point of the rest of the model.Ids are not quantities
always look up a vector, never use the id as a value.Rare tokens barely move
sensible vocab size, and tokenizer data that matches model data.Out-of-range id
check that the tokenizer and model agree on V, including special tokens.Padding has a row too
mask pad positions in attention and in the loss.Scale matters
follow your architecture's convention, then keep it consistent.