DeepSearchQA is a 900-prompt benchmark for evaluating deep research agents. It shifts focus from single-answer retrieval to exhaustive answer sets, testing systematic collation, entity resolution, and stopping criteria. Current SOTA models still face a recall-precision gap. Source: January 30 2026 DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents Google DeepMind, Google Search, Kaggle, Google Research Nikita Gupta, Riju Chatterjee, Lukas Haas, Connie Tao, Andrew Wang, Chang Liu, Hidekazu Oiwa, Elena Gribovskaya, Jan Ackermann, John Blitzer, Sasha Goldshtein, Dipanjan Das https://arxiv.org/pdf/2601.20975