This episode explores the paper Discretizing Reward Models and its argument that smooth decimal reward scores can be misleading for reinforcement learning alignment, because policies learn to exploit tiny, often meaningless differences instead of genuine quality. It explains why reward models are used for fuzzy goals like helpfulness and honesty, then digs into reward hacking, equivalence classes of equally valid answers, and the distinction between a model’s ability to separate good from bad responses versus its tendency to invent rankings among ties. The discussion also covers benchmarks such as the Ties setting and the paper’s core proposal: replacing continuous scores with a small number of ordinal reward buckets built from uncertainty estimates, pairwise equivalence judgments, and hierarchical clustering. Listeners would find it interesting because it connects an abstract modeling choice to a practical alignment problem facing modern language-model training, while also examining why the field currently seems more convinced by the diagnosis than by large-scale adoption of this exact fix.
Sources:
1. Discretizing Reward Models — Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao, 2026
http://arxiv.org/abs/2606.21795
2. Deep Reinforcement Learning from Human Preferences — Paul Christiano, Jan Leike, Tom B. Brown, Shane Legg, Dario Amodei, 2017
https://scholar.google.com/scholar?q=Deep+Reinforcement+Learning+from+Human+Preferences
3. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback — Stephen Casper, Xander Davies, Claudia Shi, Jeremy Scheurer, Dylan Hadfield-Menell, et al., 2023
https://scholar.google.com/scholar?q=Open+Problems+and+Fundamental+Limitations+of+Reinforcement+Learning+from+Human+Feedback
4. RewardBench 2: Advancing Reward Model Evaluation — Saumya Malik, Valentina Pyatkin, Sander Land, Nathan Lambert, Noah A. Smith, Hannaneh Hajishirzi, 2025
https://scholar.google.com/scholar?q=RewardBench+2%3A+Advancing+Reward+Model+Evaluation
5. Discretizing Reward Models — Vijay Viswanathan, Shiqi Wang, Devamanyu Hazarika, Chirag Nagpal, Tongshuang Wu, Graham Neubig, Yuning Mao, 2026
https://scholar.google.com/scholar?q=Discretizing+Reward+Models
6. How to Evaluate Reward Models for RLHF — Evan Frick, Tianle Li, Connor Chen, Wei-Lin Chiang, Anastasios Nikolas Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph Gonzalez, Ion Stoica, 2024
https://scholar.google.com/scholar?q=How+to+Evaluate+Reward+Models+for+RLHF
7. What Makes a Reward Model a Good Teacher? An Optimization Perspective — Noam Razin, Zixuan Wang, Hubert Strauss, Stanley Wei, Jason D. Lee, Sanjeev Arora, 2025
https://scholar.google.com/scholar?q=What+Makes+a+Reward+Model+a+Good+Teacher%3F+An+Optimization+Perspective
8. The Accuracy Paradox in RLHF: When Better Reward Models Don't Yield Better Language Models — Yanjun Chen, Dawei Zhu, Yirong Sun, Xinghao Chen, Wei Zhang, Xiaoyu Shen, 2024
https://scholar.google.com/scholar?q=The+Accuracy+Paradox+in+RLHF%3A+When+Better+Reward+Models+Don%27t+Yield+Better+Language+Models
9. Validating LLM-as-a-Judge Systems under Rating Indeterminacy — Luke Guerdan, Solon Barocas, Kenneth Holstein, Hanna Wallach, Zhiwei Steven Wu, Alexandra Chouldechova, 2025
https://scholar.google.com/scholar?q=Validating+LLM-as-a-Judge+Systems+under+Rating+Indeterminacy
10. Beyond Binary Preferences: A Principled Framework for Reward Modeling with Ordinal Feedback — Amirhossein Afsharrad, Ruida Zhou, Luca Viano, Sanjay Lall, Mohammad Ghavamzadeh, 2026
https://scholar.google.com/scholar?q=Beyond+Binary+Preferences%3A+A+Principled+Framework+for+Reward+Modeling+with+Ordinal+Feedback
11. Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts — Haoxiang Wang, Wei Xiong, Tengyang Xie, Han Zhao, Tong Zhang, 2024
https://scholar.google.com/scholar?q=Interpretable+Preferences+via+Multi-Objective+Reward+Modeling+and+Mixture-of-Experts
12. AI Post Transformers: Split Personality Training Reveals Latent Knowledge — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-08-split-personality-training-reveals-laten-c84616.mp3
13. AI Post Transformers: Robots Need More Than VLAs and World Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-10-robots-need-more-than-vlas-and-world-mod-cdab8b.mp3