modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 07

The modern decoder block

Put every upgrade from today into one block and you have the layer that Llama 3, Mistral and Qwen stack 28 to 80 times.

Before and after

The full diff

Apply the upgrades one at a time

Inside the finished block

Shapes through one Llama 3 8B block

Where one block's 218M parameters live

Four real configs

The whole block in code

Where compute goes, per generated token

Porting checklist

Exceptions and sharp edges

Qwen keeps QKV biases

do not assume bias=False everywhere.

Mistral's sliding window

pass the window to the attention kernel and the cache.

Gemma's extras

read the model's reference code, not just the config.

Tied embeddings in small models

check tie_word_embeddings.

Initialisation for training from scratch

small std (about 0.02) and scaled residual projections.