Seeing the square
Add tokens, watch the grid explode
Doubling is quadrupling
Compute and memory
When does attention dominate?
Memory to store the full score matrix
How real systems cope
FlashAttention: never build the whole square
The online softmax trick behind it
Or compute fewer scores
Exceptions and sharp edges
Short prompts: blame the MLP
profile before optimising.FlashAttention is exact
use it by default, it is not an approximation.Padding wastes squares
sort by length, pack sequences, or use variable-length kernels.Decode is linear per token
see chapter 11.Linear attention trades quality
hybrids mix a few full-attention layers back in.