modern arch

  1. 00The mapFrom the 2017 block to a 2024 LLM block
  2. 01RoPE, properlyFrequencies, caches, conventions and long context
  3. 02RMSNormRescale, skip the re-centering
  4. 03SwiGLUA gated MLP with a smooth switch
  5. 04Grouped-query attentionMany queries, few keys and values
  6. 05SDPAOne call, three backends
  7. 06FlashAttentionExact attention that respects memory
  8. 07The modern decoder blockEvery upgrade, assembled

chapter 05

SDPA: one call, three kernels

torch.nn.functional.scaled_dot_product_attention computes exactly the softmax(QKᵀ/√d)V from Day 1.

The call

Six lines become one

The arguments, drawn

Which kernel runs

The dispatcher's decision

What the math fallback costs

Masks

Boolean masks: True means attend

is_causal when L ≠ S

Shapes and control

Getting into [B, h, T, d] and back

Forcing and checking a backend

Exceptions and sharp edges

attn_mask and is_causal together

fold causal into your explicit mask, or use only is_causal.

dropout_p at eval time

dropout_p = self.p if self.training else 0.0.

Rows with nothing to see

make sure every query can see at least one key.

fp32 skips Flash

use autocast or cast the model to bf16.

The default scale

check the reference implementation's scale.