attention

  1. 00The mapEvery station of the story on one page
  2. 01Byte-level BPEText becomes integer ids
  3. 02Token embeddingsIds become vectors
  4. 03Positional informationOrder is invisible until you add it
  5. 04Sinusoidal positionsA ruler made of waves
  6. 05RoPERotate the query, rotate the key
  7. 06Q, K and VAsk, advertise, hand over
  8. 07Scaled dot-productScore, scale, softmax, mix
  9. 08Causal maskingNo peeking at the answer
  10. 09Multi-head attentionMany small lookups at once
  11. 10The n × n matrixWhere the quadratic cost lives
  12. 11KV cacheKeep the keys, skip the redo
  13. 12The decoder blockAttention and MLP on a residual highway
  14. 13Optimization and reproducibilityTrain it, then train it again identically

chapter 01

Byte-level BPE

A model cannot read letters.

Text is already numbers

Characters become bytes

Four ways to cut the same word

Before merging: draw walls

Pre-tokenization chunks

Training: learn the merges

Watch BPE learn on a tiny corpus

Encoding new text with the learned merges

How big should the vocabulary be?

The vocabulary size trade-off

Perplexity vs bits per byte

bits/byte = loss per token (nats) ÷ ln 2 ÷ bytes per token  ·  perplexity = eloss

Exceptions and sharp edges

Emoji split mid-character

when streaming, buffer bytes until they form complete UTF-8.

“ the” and “the” differ

never strip or add spaces around prompts casually.

Numbers split unevenly

split digits one by one, or in fixed groups of three.

Some languages cost more

train merges on a balanced multilingual corpus.

Case changes everything

nothing to fix, the model learns the link from data.

Ties in pair counts

break ties deterministically, for example by byte order.

Special tokens are not text

add special tokens outside the merge process.

Under-trained tokens

train the tokenizer on data like the model's.