Enterprises increasingly deploy AI agents to pull structured data from documents like invoices, contracts, and forms --- but until now, no benchmark scored accuracy, completeness, grounding, and cost together. ExtractBench fills this gap with nearly 5,000 pages across 370 documents spanning 8 business domains and 67 document types. It reveals a key tradeoff: commercial vision-language models are fast but truncate long record lists, while coding agents are accurate but expensive. This benchmark is directly useful for companies choosing extraction tools for finance, legal, or logistics workflows, helping them balance accuracy against operational cost at scale.
Authors: Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo
Paper: https://arxiv.org/abs/2607.29677v1