modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 02

RMSNorm

LayerNorm does two things: it re-centres a vector to mean zero and rescales it to unit size.

The computation

LayerNorm(x) = (x − mean) / √(var + ε) · γ + β     RMSNorm(x) = x / √(mean(x²) + ε) · γ

Side by side, one step at a time

What the two norms do geometrically

What each norm ignores

The learned part

The gain γ: one learned number per channel

What you save

Where the norms sit

Pre-norm, sandwich norm, QK-norm

What ε does

A reference implementation

Exceptions and sharp edges

Gemma stores γ − 1

check how the weight is applied before porting.

Squaring in bf16

x.float() before pow, cast back after.

Weight decay on γ

exclude norm weights and biases from weight decay.

ε inside or outside the root

match the reference implementation exactly.

Outlier channels

keep norms in higher precision; handle outliers in int8 schemes.