attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 05

RoPE, an introduction

Rotary position embedding does not add anything to the token.

Rotate, don't add

One pair, one rotation

[q′₀, q′₁] = [q₀ cos mθ − q₁ sin mθ, q₀ sin mθ + q₁ cos mθ]

Every pair has its own speed

θi = base−2i/d   with base = 10000 in the original paper

The trick: only the gap survives

Shift both, the score stays put

⟨R(m)q, R(n)k⟩ = ⟨q, R(n − m)k⟩

Where RoPE sits in attention

The rotation as a matrix

Far-apart tokens get a weaker pull

Beyond the training length

Position interpolation

Learned vs sinusoidal vs RoPE

Exceptions and sharp edges

V is never rotated

if outputs drift with absolute position, check you did not rotate v.

Which dims form a pair?

match the convention, or permute the q/k weights.

Rotated keys in the cache

pass the cache length as the position offset.

Partial rotary

read the config's rotary fraction before porting.

Precision of the angles

precompute cos and sin tables in fp32.