Without and with a cache
No cache: redo everything, every step
With a cache: compute one row, append, look up the rest
Work saved
Why K and V, but not Q
Old queries are never asked again
The cost: memory
KV cache size calculator
bytes = 2 × layers × kv_heads × head_dim × tokens × batch × bytes_per_value
Sharing K and V shrinks the cache
Prefill, then decode
Exceptions and sharp edges
Editing an earlier token
truncate the cache at the edit and recompute from there.Position offset
pass cache length as the position offset.Ragged batches
paged attention stores the cache in fixed-size blocks.Evicting old entries
keep the first few tokens plus a recent window.Quantised caches
measure on long-context tasks before shipping.Shared prefixes
prompt caching reuses the prefix's KV blocks.