Single head vs multi-head
One softmax, two needs
Split the width, do not multiply it
What different heads learn
Four heads, four habits
The same four heads as matrices
Putting heads back together
Concatenate, then mix with WO
MultiHead(x) = concat(head₁, …, headh) · WO
More heads, same parameter bill
Heads can share keys and values
Exceptions and sharp edges
dmodel must divide by h
pick h so d_model / h is a whole number, usually 64 or 128.Heads that are too narrow
keep d_head around 64 to 128 and scale heads with width.Many heads are redundant
useful for compression, but you still need them during training.Attention sinks
keep the first tokens in the cache when trimming long contexts.The reshape bug
check one head's output against a single-head reference.