Vision-language model agents increasingly need to act on 3D scenes, not just describe them, yet existing benchmarks mostly score text responses or single-object tasks. SceneActBench introduces a unified agent-environment loop evaluating five 3D action tasks across 210 source instances and 520 task cases, using images, video frames, and 3D assets, scored against hidden ground truth with geometric metrics. Testing eleven proprietary VLM configurations revealed inconsistent performance (38.6-50.2 overall) with no model excelling across all tasks. Applications include benchmarking and improving embodied AI, robotics manipulation planning, and 3D-scene-grounded agent development for real-world spatial reasoning tasks.
Authors: Yifei Zhao, Xiangxin Zhou, Wenhao Yang, Jiaqi Tang, Pu Jian, Huanjin Yao, Jiarui Yao, Haowei Lin, Chunchao Guo, Zhuo Chen, Wenkai Lyu, Jianzhu Ma, Xueqian Wang, Wenxi Zhu
Paper: https://arxiv.org/abs/2607.22393v1