This episode explores a 2025 study on fine-tuning large language models to predict how people respond in social science experiments, asking whether trained models can simulate new studies more reliably than prompting alone. It explains how the researchers built SOCSCI210, a dataset of 2.9 million responses from more than 400,000 participants across 210 TESS experiments, and why standardizing those studies into respondent-condition-question-answer records is central to the method. The discussion breaks down the paper’s evaluation criteria, including out-of-distribution generalization, distribution matching via Wasserstein distance, normalized individual accuracy, and treatment-effect recovery, to show the difference between sounding plausible and preserving real experimental patterns. Listeners would find it interesting because it treats LLMs not as chatbots but as possible “wind tunnels” for testing study designs in advance, while also confronting the risk that a convincing simulator could still get causal effects wrong.
Sources:
1. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025
http://arxiv.org/abs/2509.05830
2. Out of One, Many: Using Language Models to Simulate Human Samples — Lisa P. Argyle, Ethan C. Busby, Nancy Fulda, Joshua Gubler, Christopher Rytting, David Wingate, 2022
https://scholar.google.com/scholar?q=Out+of+One%2C+Many%3A+Using+Language+Models+to+Simulate+Human+Samples
3. LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals — Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Michael S. Bernstein, et al., 2024
https://scholar.google.com/scholar?q=LLM+Agents+Grounded+in+Self-Reports+Enable+General-Purpose+Simulation+of+Individuals
4. Large Language Models Show Human-like Social Desirability Biases in Survey Responses — Aadesh Salecha, Molly E. Ireland, Shashanka Subrahmanya, Joao Sedoc, Lyle H. Ungar, Johannes C. Eichstaedt, 2024
https://scholar.google.com/scholar?q=Large+Language+Models+Show+Human-like+Social+Desirability+Biases+in+Survey+Responses
5. Finetuning LLMs for Human Behavior Prediction in Social Science Experiments — Akaash Kolluri, Shengguang Wu, Joon Sung Park, Michael S. Bernstein, 2025
https://scholar.google.com/scholar?q=Finetuning+LLMs+for+Human+Behavior+Prediction+in+Social+Science+Experiments
6. Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies — Gati Aher, Rosa I. Arriaga, Adam Tauman Kalai, 2022
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Simulate+Multiple+Humans+and+Replicate+Human+Subject+Studies
7. Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management — Ziyan Cui, Ning Li, Huaikang Zhou, 2024
https://scholar.google.com/scholar?q=Can+Large+Language+Models+Replace+Human+Subjects%3F+A+Large-Scale+Replication+of+Scenario-Based+Experiments+in+Psychology+and+Management
8. Using Large Language Models to Create AI Personas for Replication, Generalization and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings — Leo Yeykelis, Kaavya Pichai, James J. Cummings, Byron Reeves, 2024
https://scholar.google.com/scholar?q=Using+Large+Language+Models+to+Create+AI+Personas+for+Replication%2C+Generalization+and+Prediction+of+Media+Effects%3A+An+Empirical+Test+of+133+Published+Experimental+Research+Findings
9. This human study did not involve human subjects: Validating LLM simulations as behavioral evidence — Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw, 2026
https://scholar.google.com/scholar?q=This+human+study+did+not+involve+human+subjects%3A+Validating+LLM+simulations+as+behavioral+evidence
10. Centaur: a Foundation Model of Human Cognition — Marcel Binz et al., 2024
https://scholar.google.com/scholar?q=Centaur%3A+a+Foundation+Model+of+Human+Cognition
11. Language Model Fine-Tuning on Scaled Survey Data for Predicting Distributions of Public Opinions — Joseph Suh, Erfan Jahanparast, Suhong Moon, Minwoo Kang, Serina Chang, 2025
https://scholar.google.com/scholar?q=Language+Model+Fine-Tuning+on+Scaled+Survey+Data+for+Predicting+Distributions+of+Public+Opinions
12. Beyond Demographics: Fine-tuning Large Language Models to Predict Individuals' Subjective Text Perceptions — Matthias Orlikowski, Jiaxin Pei, Paul Rottger, Philipp Cimiano, David Jurgens, Dirk Hovy, 2025
https://scholar.google.com/scholar?q=Beyond+Demographics%3A+Fine-tuning+Large+Language+Models+to+Predict+Individuals%27+Subjective+Text+Perceptions
13. Large Language Models that Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups — Angelina Wang, Jamie Morgenstern, John P. Dickerson, 2025
https://scholar.google.com/scholar?q=Large+Language+Models+that+Replace+Human+Participants+Can+Harmfully+Misportray+and+Flatten+Identity+Groups
14. Beyond Believability: Accurate Human Behavior Simulation with Fine-Tuned LLMs — Yuxuan Lu et al., 2025
https://scholar.google.com/scholar?q=Beyond+Believability%3A+Accurate+Human+Behavior+Simulation+with+Fine-Tuned+LLMs
15. The Prompt Makes the Person(a): A Systematic Evaluation of Sociodemographic Persona Prompting for Large Language Models — Marlene Lutz et al., 2025
https://arxiv.org/abs/2507.16076
16. Prompt Fairness: Sub-group Disparities in LLMs — Meiyu Zhong, Noel Teku, Ravi Tandon, 2025
https://arxiv.org/abs/2511.19956
17. Can Persona-Prompted LLMs Emulate Subgroup Values? An Empirical Analysis of Generalisability and Fairness in Cultural Alignment — Bryan Chen Zhengyu Tan et al., 2026
https://arxiv.org/abs/2604.12851
18. Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent-Based Persona Routing with PRISM — Zizhao Hu, Mohammad Rostami, Jesse Thomason, 2026
https://arxiv.org/abs/2603.18507
19. Causality for Large Language Models — Anpeng Wu et al., 2024
https://arxiv.org/abs/2410.15319
20. AI Post Transformers: PaperBench: Can AI Replicate AI Research? — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-06-17-paperbench-can-ai-replicate-ai-research-862944.mp3
21. AI Post Transformers: When Many-Shot CoT Becomes Test-Time Learning — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-05-15-when-many-shot-cot-becomes-test-time-lea-c25bfe.mp3
Interactive Visualization: Fine-Tuning LLMs for Human Behavior Prediction