Before and after
The full diff
Apply the upgrades one at a time
Inside the finished block
Shapes through one Llama 3 8B block
Where one block's 218M parameters live
Four real configs
The whole block in code
Where compute goes, per generated token
Porting checklist
Exceptions and sharp edges
Qwen keeps QKV biases
do not assume bias=False everywhere.Mistral's sliding window
pass the window to the attention kernel and the cache.Gemma's extras
read the model's reference code, not just the config.Tied embeddings in small models
check tie_word_embeddings.Initialisation for training from scratch
small std (about 0.02) and scaled residual projections.