LLM agents are increasingly used to automate scientific hypothesis testing, but they often make subtle statistical errors leading to invalid conclusions, even when their code execution is correct—a failure mode not captured by existing benchmarks. This paper introduces P-Bench, a benchmark of 425 hypothesis-testing tasks spanning economics, biology, and medicine, and Fisher-R1, an LLM agent trained via reinforcement learning for rigorous statistical reasoning. Applications include automated scientific research assistants, data analysis pipelines, and tools supporting empirical claims in academic or industry settings, where Fisher-R1 substantially outperformed strong baselines like GPT-5.4 and DeepSeek-V4-Pro on statistical validity.
Paper: https://arxiv.org/abs/2608.07437