Text is already numbers
Characters become bytes
Four ways to cut the same word
Before merging: draw walls
Pre-tokenization chunks
Training: learn the merges
Watch BPE learn on a tiny corpus
Encoding new text with the learned merges
How big should the vocabulary be?
The vocabulary size trade-off
Perplexity vs bits per byte
bits/byte = loss per token (nats) ÷ ln 2 ÷ bytes per token · perplexity = eloss
Exceptions and sharp edges
Emoji split mid-character
when streaming, buffer bytes until they form complete UTF-8.“ the” and “the” differ
never strip or add spaces around prompts casually.Numbers split unevenly
split digits one by one, or in fixed groups of three.Some languages cost more
train merges on a balanced multilingual corpus.Case changes everything
nothing to fix, the model learns the link from data.Ties in pair counts
break ties deterministically, for example by byte order.Special tokens are not text
add special tokens outside the merge process.Under-trained tokens
train the tokenizer on data like the model's.