AI Post Transformers

When LLM Judges Become Coin Flips


Listen Later

This episode explores a March 2026 paper arguing that LLM-based judges are an unreliable way to measure jailbreak success and adversarial robustness. It explains how modern safety evaluations rely on judge models to score harmful outputs, then walks through why those judges can break under attack shift, model shift, and data shift, sometimes degrading to near coin-flip reliability. The discussion connects this critique to benchmarks such as MT-Bench, HarmBench, and StrongREJECT, and examines how weaknesses in the judging pipeline can inflate or distort reported attack success rates. Listeners would find it interesting because it challenges whether many headline jailbreak results are exposing real model failures or simply failures in the grading system.
Sources:
1. A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial Robustness — Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, Gauthier Gidel, Stephan Günnemann, 2026
http://arxiv.org/abs/2603.06594
2. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena — Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zi Lin, Zhuohan Li, Joseph E. Gonzalez, Ion Stoica and others, 2023
https://scholar.google.com/scholar?q=Judging+LLM-as-a-Judge+with+MT-Bench+and+Chatbot+Arena
3. Large Language Models are not Fair Evaluators — Peiyi Wang, Lei Li, Liang Chen, Zefan Cai, Dawei Zhu, Binghuai Lin, Yunbo Cao, Qi Liu, Tianyu Liu, Zhifang Sui, 2023
https://scholar.google.com/scholar?q=Large+Language+Models+are+not+Fair+Evaluators
4. An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Models are Task-specific Classifiers — Hui Huang, Yingqi Qu, Jing Liu, Muyun Yang, Tiejun Zhao, 2024
https://scholar.google.com/scholar?q=An+Empirical+Study+of+LLM-as-a-Judge+for+LLM+Evaluation%3A+Fine-tuned+Judge+Models+are+Task-specific+Classifiers
5. JudgeBench: A Benchmark for Evaluating LLM-based Judges — Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y. Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, Ion Stoica, 2024
https://scholar.google.com/scholar?q=JudgeBench%3A+A+Benchmark+for+Evaluating+LLM-based+Judges
6. Universal and Transferable Adversarial Attacks on Aligned Language Models — Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, Matt Fredrikson, 2023
https://scholar.google.com/scholar?q=Universal+and+Transferable+Adversarial+Attacks+on+Aligned+Language+Models
7. HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal — Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, Dan Hendrycks, 2024
https://scholar.google.com/scholar?q=HarmBench%3A+A+Standardized+Evaluation+Framework+for+Automated+Red+Teaming+and+Robust+Refusal
8. A StrongREJECT for Empty Jailbreaks — Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, Sam Toyer, 2024
https://scholar.google.com/scholar?q=A+StrongREJECT+for+Empty+Jailbreaks
9. JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models — Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J. Pappas, Florian Tramèr, Hamed Hassani, Eric Wong, 2024
https://scholar.google.com/scholar?q=JailbreakBench%3A+An+Open+Robustness+Benchmark+for+Jailbreaking+Large+Language+Models
10. Dataset Shift in Machine Learning — Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, Neil D. Lawrence (editors), 2008
https://scholar.google.com/scholar?q=Dataset+Shift+in+Machine+Learning
11. Can You Trust Your Model's Uncertainty? Evaluating Predictive Uncertainty Under Dataset Shift — Yaniv Ovadia, Emily Fertig, Jie Ren, Zachary Nado, D. Sculley, Sebastian Nowozin, Joshua Dillon, Balaji Lakshminarayanan, Jasper Snoek, 2019
https://scholar.google.com/scholar?q=Can+You+Trust+Your+Model%27s+Uncertainty%3F+Evaluating+Predictive+Uncertainty+Under+Dataset+Shift
12. Measuring Robustness to Natural Distribution Shifts in Image Classification — Rohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini, Benjamin Recht, Ludwig Schmidt, 2020
https://scholar.google.com/scholar?q=Measuring+Robustness+to+Natural+Distribution+Shifts+in+Image+Classification
13. WILDS: A Benchmark of in-the-Wild Distribution Shifts — Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Percy Liang and others, 2021
https://scholar.google.com/scholar?q=WILDS%3A+A+Benchmark+of+in-the-Wild+Distribution+Shifts
14. LLMs Cannot Reliably Judge (Yet?): A Comprehensive Assessment on the Robustness of LLM-as-a-Judge — Shuaizhi Li, Chenxu Xu, Jiazhu Wang, Xianyu Gong, Cheng Chen, Jun Zhang, Junjie Wang, Kit Lam, and Shouling Ji, 2025
https://scholar.google.com/scholar?q=LLMs+Cannot+Reliably+Judge+%28Yet%3F%29%3A+A+Comprehensive+Assessment+on+the+Robustness+of+LLM-as-a-Judge
15. Confusion Is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs — Yixin Yan, Shichao Sun, Zhen Wang, Yifan Lin, Zeyu Duan, Zhenzhen Zheng, Mingyu Liu, Zhenfei Yin, and Jie Zhang, 2025
https://scholar.google.com/scholar?q=Confusion+Is+the+Final+Barrier%3A+Rethinking+Jailbreak+Evaluation+and+Investigating+the+Real+Misuse+Threat+of+LLMs
16. Know Thy Judge: On the Robustness Meta-Evaluation of LLM Safety Judges — Francisco Eiras et al., 2025
https://scholar.google.com/scholar?q=Know+Thy+Judge%3A+On+the+Robustness+Meta-Evaluation+of+LLM+Safety+Judges
17. Comparison Requires Valid Measurement: Rethinking Attack Success Rate Comparisons in AI Red Teaming — Alexandra Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, Hanna Wallach, 2025/2026
https://scholar.google.com/scholar?q=Comparison+Requires+Valid+Measurement%3A+Rethinking+Attack+Success+Rate+Comparisons+in+AI+Red+Teaming
18. How to Correctly Report LLM-as-a-Judge Evaluations — Chungpa Lee et al., 2025
https://scholar.google.com/scholar?q=How+to+Correctly+Report+LLM-as-a-Judge+Evaluations
19. Embracing Ambiguity: Bayesian Nonparametrics and Stakeholder Participation for Ambiguity-Aware Safety Evaluation — Yanan Long, 2025/2026
https://scholar.google.com/scholar?q=Embracing+Ambiguity%3A+Bayesian+Nonparametrics+and+Stakeholder+Participation+for+Ambiguity-Aware+Safety+Evaluation
20. AI Post Transformers: Multidimensional Safety Evaluation of Frontier AI Models — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/multidimensional-safety-evaluation-of-frontier-ai-models/
21. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/
22. AI Post Transformers: AI Agent Traps and Prompt Injection — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-02-ai-agent-traps-and-prompt-injection-7ce4ba.mp3
Interactive Visualization: When LLM Judges Become Coin Flips
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof