Inside one block
One trip through a pre-norm block
h = x + Attn(Norm(x)) out = h + MLP(Norm(h))
The residual stream is a shared notebook
Pre-norm vs post-norm
The pieces
LayerNorm and RMSNorm
The MLP: widen, bend, narrow
Attention mixes across tokens, the MLP within a token
The whole model
From the last block to the next word
Where the parameters live
Exceptions and sharp edges
Forgetting the final norm
norm → unembed, always.The stream grows with depth
scale residual-branch init by 1/√(2N), as GPT-2 does.SwiGLU has three matrices
read hidden size from the config, never assume 4d.Norm epsilon
compute norms in fp32, keep ε around 1e-5 or 1e-6.Dropout placement
most large LLMs use no dropout in pretraining at all.