attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 03

Positional information

Attention compares every token with every other token, and it does not care where they sit.

The blind spot

The shuffle test

Two kinds of “where”

Learned absolute positions

Token vector + position vector

xi = E[idi] + P[i]

The position table

What the learned table ends up looking like

Where learned positions break

The maximum-length wall

Same relation, different numbers

Three answers, three injection points

Exceptions and sharp edges

Left padding shifts positions

compute position ids from the attention mask, not from the raw index.

Packed documents

pick one, and use matching attention masks.

Under-trained tail rows

include enough long examples, or use a scheme without a table.

No positions is not zero information

a curiosity, not a replacement for real positions.

Swapping position scheme

decide at the start, or plan a fine-tuning phase.