May 13, 2026•arxiv.org•▲distillation▲ml▲researchA Survey of On-Policy Distillation for Large Language ModelsLiterature review outlining sequence-level distillation techniques across MiniLLM, GKD, and f-divergence generalizations. Analyzes teacher guidance strategies during autoregressive sampling.
May 10, 2026•nrehiew.github.io•▲distillation▲rl▲mlSFT, RL, and On-Policy Distillation Through a Distributional LensA unified mathematical formulation connecting supervised fine-tuning (forward KL), RLHF (reverse KL with reward shaping), and on-policy distillation. Clarifies why on-policy student generations prevent mode collapse and exposure bias.