Language models are typically trained on plain text extracted from documents, discarding layout, figures, and typography that carry meaningful information. This paper investigates training foundation models directly on visual renderings of documents instead of extracted text, testing whether "seeing" a page beats "reading" it. Across multiple model backbones and benchmarks, visual pretraining outperforms text-only pretraining on the same source material. Applications include more efficient and capable foundation models for tasks involving richly formatted documents, web pages, scientific papers, and forms - domains where converting to plain text currently discards valuable structural and visual information.
Authors: Yiming Zhang, Zhonghan Zhao, Wenwei Zhang, Haiteng Zhao, Tianyang Lin, Yunhua Zhou, Demin Song, Kuikun Liu, Haochen Ye, Haian Huang, Yuzhe Gu, Haijun Lv, Qipeng Guo, Bin Liu, Gaoang Wang, Kai Chen
Paper: https://arxiv.org/abs/2607.09657v1