modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 01

RoPE, properly

Day 1 introduced the idea: rotate queries and keys by their position, and the dot product sees only the distance.

Where it lives

RoPE inside one attention layer, with shapes

The frequency table

One frequency per pair, set by the base

inv_freq = 1 / base ** (arange(0, d_h, 2) / d_h)     wavelengthi = 2π / θi

Build cos and sin once, look them up forever

Pairing and rotate_half

Two ways to choose the pairs

rotate_half, step by step

rotate_half([x₁, x₂]) = [−x₂, x₁]     q′ = q · cos + rotate_half(q) · sin

Positions at decode time

The new token starts where the cache ends

position_ids = cache_len + arange(new_tokens)

Proof by picture: scores depend only on distance

Stretching to longer contexts

Interpolation, NTK-aware, YaRN

Exceptions and sharp edges

Rotating v

rotate exactly q and k.

Wrong base at inference

read rope_theta from the config, never hardcode it.

Converted checkpoints

use the conversion script that ships with the model.

Left padding

derive position_ids from the attention mask.

Low-precision tables

build inv_freq and angles in fp32, cast cos and sin at the end.

Partial rotary

check for a rotary fraction or rotary_dim in the config.