octlm · day 1 · transformer foundations
Attention, drawn one stroke at a time
Thirteen short chapters that follow one sentence from raw bytes to a predicted next token.
The whole trip, one sentence long
Our running example
How to read the drawings
Shapes along the way
Chapters
chapter 01 · act 1Byte-level BPEText becomes integer idschapter 02 · act 1Token embeddingsIds become vectorschapter 03 · act 2Positional informationOrder is invisible until you add itchapter 04 · act 2Sinusoidal positionsA ruler made of waveschapter 05 · act 2RoPERotate the query, rotate the keychapter 06 · act 3Q, K and VAsk, advertise, hand overchapter 07 · act 3Scaled dot-productScore, scale, softmax, mixchapter 08 · act 3Causal maskingNo peeking at the answerchapter 09 · act 3Multi-head attentionMany small lookups at oncechapter 10 · act 3The n × n matrixWhere the quadratic cost liveschapter 11 · act 3KV cacheKeep the keys, skip the redochapter 12 · act 4The decoder blockAttention and MLP on a residual highwaychapter 13 · act 4Optimization and reproducibilityTrain it, then train it again identically