The loop
One training step, six moves
Cross-entropy: the price of surprise
loss = − log p(target) · perplexity = emean loss
Stepping downhill
Learning rate: too small, just right, too big
SGD vs Adam in a narrow valley
Warmup, then cosine decay
Gradient clipping
g ← g · min(1, c / ‖g‖)
Reading loss curves
Reproducibility
Same seed, different seed, nondeterministic kernels
The reproducibility checklist
Exceptions and sharp edges
Atomic adds on GPU
torch.use_deterministic_algorithms(True), at some speed cost.Different hardware
record versions; expect “statistically same”, not “bitwise same”, across machines.Data loader workers
seed each worker from the base seed and worker id.Resuming loses state
checkpoint optimizer, scheduler, RNG states and the data position.Changing batch size
retune LR, or scale it with batch size and re-check.Loss spikes
clip gradients, lower peak LR, or rewind and skip the batch.