A first-principles-to-frontier survey of how reinforcement learning is adapted to language models, explaining the progression from classical RL concepts through RLHF and newer reasoning-oriented training methods.
- Frames an LLM as a policy that generates token sequences, with rewards assigned to outputs or trajectories rather than individual actions.
- Decomposes the RLHF pipeline into supervised fine-tuning, preference-based reward modeling, and policy optimization, including the role of KL regularization in preventing drift from the reference model.
- Connects modern approaches such as verifiable rewards and group-relative policy updates to the practical challenge of improving reasoning without relying exclusively on human preference labels.