Building robots that can understand and interact with the physical world requires massive amounts of 3D training data — but capturing that data with multi-camera rigs is expensive and impractical at scale. HAT-4D proposes using ordinary monocular video as a data source, reconstructing the 3D geometry and temporal dynamics of multiple interacting objects with the help of vision-language models and targeted human feedback. It introduces a new benchmark, MVOIK-4D, for evaluating such reconstructions on physical plausibility and temporal consistency. Applications include scalable embodied AI training data pipelines, robotic manipulation learning, AR/VR scene reconstruction, and virtual environment generation from real-world video footage.
Authors: Jiaxin Li, Yuxiang Wu, Zhenkai Zhang, Xinrui Shi, Haoyuan Wang, Yichen Zhao, Su Linxiang, Chenyang Yu, Mingyu Zhang, Yifan Ding, Boran Wen, Li Zhang, Ruiyang Liu, Yong-Lu Li
Paper: https://arxiv.org/abs/2606.28215v1