The computation
LayerNorm(x) = (x − mean) / √(var + ε) · γ + β RMSNorm(x) = x / √(mean(x²) + ε) · γ
Side by side, one step at a time
What the two norms do geometrically
What each norm ignores
The learned part
The gain γ: one learned number per channel
What you save
Where the norms sit
Pre-norm, sandwich norm, QK-norm
What ε does
A reference implementation
Exceptions and sharp edges
Gemma stores γ − 1
check how the weight is applied before porting.Squaring in bf16
x.float() before pow, cast back after.Weight decay on γ
exclude norm weights and biases from weight decay.ε inside or outside the root
match the reference implementation exactly.Outlier channels
keep norms in higher precision; handle outliers in int8 schemes.