Introduction
Why do you need reasoning models?
()
1. The Power of Reasoning Models
The Shift to Reasoning Models
()
The reasoning landscape
()
2. The Four Approaches to Building Reasoning Models
Inference-time scaling
()
Pure reinforcement learning (RL)
()
Supervised fine-tuning (SFT) and RL
()
Distillation and pure SFT
()
3. Test-Time Compute Scaling Deep Dive
Majority voting and self-consistency
()
Best-of-n and weighted aggregation
()
Beam search with process reward models
()
Diverse Verifier Tree Search (DVTS)
()
4. Reinforcement Learning for Reasoning
Beyond RLHF: Group Relative Policy Optimization (GRPO)
()
Reward functions for reasoning
()
The "aha moment": Self-verification through RL
()
5. Building Efficient Reasoning Systems
Compute-optimal scaling in production
()
Budget-friendly reasoning models
()
Balancing cost and performance
()
Conclusion
Future directions in reasoning LLMs
()