← overview
Deep dive A · post-training · source snapshots 21 and 22 September 2026
How production recipes sequence SFT, preference tuning, RL and on-policy distillation
Post-training is no longer one stage: documented pipelines chain SFT, preference optimization, verifiable-reward RL and, increasingly, on-policy distillation, in orders that differ by lab. This page records each recipe in the source's own words, the frameworks per sub-stage, the stabilizers production runs rely on, and which audited tasks reach the stage.