attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 08

Causal masking

A language model learns by predicting the next token.

Why hide the future

The cheating problem

Seven lessons from one forward pass

The mask itself

Building the triangle row by row

mask[i, j] = 0 if j ≤ i,   −∞ if j > i     scores ← scores + mask

Mask before softmax, not after

A live causality test

More masks

Causal + padding = combined mask

Three mask shapes, three model families

Exceptions and sharp edges

Off by one on the diagonal

use torch.triu(…, diagonal=1) for the masked part.

Suspiciously low loss

run the causality test from fig 8.5 on your model.

Pad rows with nothing to see

let pad rows see something, and mask them out of the loss.

Generation needs no triangle

pass the right mask for prefill vs decode, see chapter 11.

Sliding windows

combine masks by AND, never by overwriting.