AI Post Transformers

PALOMA: Benchmarking Language Model Fit Across Domains


Listen Later

This episode explores PALOMA, a NeurIPS 2024 benchmark designed to measure how well language models fit many different language distributions instead of relying on a single average perplexity score. It explains why one global loss number can hide important weaknesses across domains such as specific subreddits, scientific writing, or programming languages, and highlights PALOMA’s fine-grained setup across 546 English and code domains from 16 sources. The discussion places PALOMA in context with earlier language-model evaluation traditions, scaling-law work, and broader benchmark efforts like HELM, while arguing that evaluation design determines what claims researchers can actually make. Listeners would find it interesting for its clear case that better measurement, data curation, and decontamination can reveal model behavior that broad headline metrics often miss.
Sources:
1. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, Jesse Dodge, 2023
http://arxiv.org/abs/2312.10523
2. One Billion Word Benchmark for Measuring Progress in Statistical Language Modeling — Ciprian Chelba, Tomas Mikolov, Mike Schuster, Qi Ge, Thorsten Brants, et al., 2013
https://scholar.google.com/scholar?q=One+Billion+Word+Benchmark+for+Measuring+Progress+in+Statistical+Language+Modeling
3. Scaling Laws for Neural Language Models — Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Dario Amodei, et al., 2020
https://scholar.google.com/scholar?q=Scaling+Laws+for+Neural+Language+Models
4. Training Compute-Optimal Large Language Models — Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Jack W. Rae, Oriol Vinyals, Laurent Sifre, et al., 2022
https://scholar.google.com/scholar?q=Training+Compute-Optimal+Large+Language+Models
5. Paloma: A Benchmark for Evaluating Language Model Fit — Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Kyle Richardson, Jesse Dodge, et al., 2024
https://scholar.google.com/scholar?q=Paloma%3A+A+Benchmark+for+Evaluating+Language+Model+Fit
6. M2D2: A Massively Multi-domain Language Modeling Dataset — Machel Reid, Victor Zhong, Suchin Gururangan, Luke Zettlemoyer, 2022
https://scholar.google.com/scholar?q=M2D2%3A+A+Massively+Multi-domain+Language+Modeling+Dataset
7. Holistic Evaluation of Language Models — Percy Liang et al., 2022
https://scholar.google.com/scholar?q=Holistic+Evaluation+of+Language+Models
8. Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models — Hong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu Ma, 2022
https://scholar.google.com/scholar?q=Same+Pre-training+Loss%2C+Better+Downstream%3A+Implicit+Bias+Matters+for+Language+Models
9. Language Model Evaluation Beyond Perplexity — Clara Meister, Ryan Cotterell, 2021
https://scholar.google.com/scholar?q=Language+Model+Evaluation+Beyond+Perplexity
10. Unsupervised Domain Clusters in Pretrained Language Models — Roee Aharoni, Yoav Goldberg, 2020
https://scholar.google.com/scholar?q=Unsupervised+Domain+Clusters+in+Pretrained+Language+Models
11. DataComp-LM: In search of the next generation of training sets for language models — Jeffrey Li et al., 2024
https://scholar.google.com/scholar?q=DataComp-LM%3A+In+search+of+the+next+generation+of+training+sets+for+language+models
12. Rethinking Perplexity: Revealing the Impact of Input Length on Perplexity Evaluation in LLMs — Letian Cheng et al., 2026
https://scholar.google.com/scholar?q=Rethinking+Perplexity%3A+Revealing+the+Impact+of+Input+Length+on+Perplexity+Evaluation+in+LLMs
13. Rethinking Benchmark and Contamination for Language Models with Rephrased Samples — Shuo Yang et al., 2023
https://scholar.google.com/scholar?q=Rethinking+Benchmark+and+Contamination+for+Language+Models+with+Rephrased+Samples
14. PaCoST: Paired Confidence Significance Testing for Benchmark Contamination Detection in Large Language Models — Huixuan Zhang et al., 2024
https://scholar.google.com/scholar?q=PaCoST%3A+Paired+Confidence+Significance+Testing+for+Benchmark+Contamination+Detection+in+Large+Language+Models
15. RegMix: Data Mixture as Regression for Language Model Pre-training — Qian Liu et al., 2024
https://scholar.google.com/scholar?q=RegMix%3A+Data+Mixture+as+Regression+for+Language+Model+Pre-training
16. Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance — Jiasheng Ye et al., 2024
https://scholar.google.com/scholar?q=Data+Mixing+Laws%3A+Optimizing+Data+Mixtures+by+Predicting+Language+Modeling+Performance
17. The Fools are Certain; the Wise are Doubtful: Exploring LLM Confidence in Code Completion — Zoe Kotti et al., 2025
https://scholar.google.com/scholar?q=The+Fools+are+Certain%3B+the+Wise+are+Doubtful%3A+Exploring+LLM+Confidence+in+Code+Completion
18. AI Post Transformers: Model-Aware Tokenizer Transfer for Multilingual LLMs — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-16-model-aware-tokenizer-transfer-for-multi-90666c.mp3
19. AI Post Transformers: LLM Benchmark Robustness to Linguistic Variation — Hal Turing & Dr. Ada Shannon, 2025
https://podcast.do-not-panic.com/episodes/llm-benchmark-robustness-to-linguistic-variation/
...more
View all episodesView all episodes
Download on the App Store

AI Post TransformersBy mcgrof