Building benchmarks for robots and other embodied AI is slow, and an early mistake can quietly spoil the result. Embodied-BenchForge has agents build the benchmark, check each intermediate piece, and redo only the parts that fail. It produced six question-answering benchmarks and one interactive benchmark with 220 tasks.
Authors: Baoyang Jiang, Fengchun Zhang, Leyuan Wang, Haotian Li, Yida Wang, Zhe Ji, Jinshan Lai, Xi Ren, Danyang Li, Zheng Yang, Jianwei Hu, Qiang Ma
Paper: https://arxiv.org/abs/2609.13082v1