chapter 07
Scaled dot-product attention
One line of maths runs every attention head in every modern LLM.
Attention(Q, K, V) = softmax( Q Kᵀ / √dk + mask ) · V
The whole computation
Six steps, one matrix at a time
Every row is chapter 6 again
Why divide by √dk
Dot products spread out as width grows
Large scores freeze softmax
Softmax, carefully
Softmax step by step, with the overflow trap
softmax(s)i = esi − max(s) ÷ Σj esj − max(s) (same answer, no overflow)