This episode explores X-LLM, a 2023 system that treats images, video, and speech as foreign languages a frozen ChatGLM can learn to read through learned modality-to-language bridges. It breaks down the paper’s architecture, including Q-Former-based visual adapters and a separate speech pipeline with continuous integrate-and-fire modules, to show how three sensory routes feed a single dialogue model instead of one end-to-end multimodal transformer. The discussion argues that X-LLM mattered less as proof of a universal multimodal theory than as a practical open-model recipe shaped by 2023 compute limits, with its Chinese-language backbone playing a real methodological role rather than serving as background context. Listeners get a sharp comparison between this bridge-based approach and later end-to-end systems such as GPT-4o and Gemini 1.5, making the episode useful for understanding how modern multimodal assistants actually evolved.
Sources:
1. X-LLM: Bootstrapping Advanced Large Language Models by Treating Multi-Modalities as Foreign Languages — Feilong Chen, Minglun Han, Haozhi Zhao, Qingyang Zhang, Jing Shi, Shuang Xu, Bo Xu, 2023
http://arxiv.org/abs/2305.04160
2. Multimodal Few-Shot Learning with Frozen Language Models — Maria Tsimpoukelli, Jacob Menick, Oriol Vinyals, Felix Hill, 2021
https://scholar.google.com/scholar?q=Multimodal+Few-Shot+Learning+with+Frozen+Language+Models
3. Flamingo: a Visual Language Model for Few-Shot Learning — Jean-Baptiste Alayrac, Jeff Donahue, Karen Simonyan, Oriol Vinyals, Andrew Zisserman, 2022
https://scholar.google.com/scholar?q=Flamingo%3A+a+Visual+Language+Model+for+Few-Shot+Learning
4. BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models — Junnan Li, Dongxu Li, Silvio Savarese, Steven Hoi, 2023
https://scholar.google.com/scholar?q=BLIP-2%3A+Bootstrapping+Language-Image+Pre-training+with+Frozen+Image+Encoders+and+Large+Language+Models
5. SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities — Dong Zhang, Shimin Li, Xin Zhang, Xipeng Qiu, 2023
https://scholar.google.com/scholar?q=SpeechGPT%3A+Empowering+Large+Language+Models+with+Intrinsic+Cross-Modal+Conversational+Abilities
6. Visual Instruction Tuning — Haotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae Lee, 2023
https://scholar.google.com/scholar?q=Visual+Instruction+Tuning
7. MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models — Deyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li, Mohamed Elhoseiny, 2023
https://scholar.google.com/scholar?q=MiniGPT-4%3A+Enhancing+Vision-Language+Understanding+with+Advanced+Large+Language+Models
8. PaLM-E: An Embodied Multimodal Language Model — Danny Driess et al., 2023
https://scholar.google.com/scholar?q=PaLM-E%3A+An+Embodied+Multimodal+Language+Model
9. CIF: Continuous Integrate-and-Fire for End-to-End Speech Recognition — Linhao Dong, Bo Xu, 2019
https://scholar.google.com/scholar?q=CIF%3A+Continuous+Integrate-and-Fire+for+End-to-End+Speech+Recognition
10. VL-JEPA: Joint Embedding Predictive Architecture for Vision-language — Delong Chen et al., 2025
https://scholar.google.com/scholar?q=VL-JEPA%3A+Joint+Embedding+Predictive+Architecture+for+Vision-language
11. TokenPacker: Efficient Visual Projector for Multimodal LLM — Wentong Li et al., 2024
https://scholar.google.com/scholar?q=TokenPacker%3A+Efficient+Visual+Projector+for+Multimodal+LLM
12. Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM — Donghwan Chi et al., 2025
https://scholar.google.com/scholar?q=Slot-MLLM%3A+Object-Centric+Visual+Tokenization+for+Multimodal+LLM
13. Auto-Encoding Morph-Tokens for Multimodal LLM — Kaihang Pan et al., 2024
https://scholar.google.com/scholar?q=Auto-Encoding+Morph-Tokens+for+Multimodal+LLM
14. ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention — Wenjie Liu et al., 2026
https://scholar.google.com/scholar?q=ViCA%3A+Efficient+Multimodal+LLMs+with+Vision-Only+Cross-Attention
15. F-LMM: Grounding Frozen Large Multimodal Models — Size Wu et al., 2024
https://scholar.google.com/scholar?q=F-LMM%3A+Grounding+Frozen+Large+Multimodal+Models
16. MultiModal-GPT: A Vision and Language Model for Dialogue with Humans — Tao Gong et al., 2023
https://scholar.google.com/scholar?q=MultiModal-GPT%3A+A+Vision+and+Language+Model+for+Dialogue+with+Humans
17. AI Post Transformers: UniVideo: Unified Video Understanding, Generation, and Editing — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/univideo-unified-video-understanding-generation-and-editing/
18. AI Post Transformers: DeepSeek-OCR: Contexts Optical Compression — Hal Turing & Dr. Ada Shannon, Sat,
https://podcast.do-not-panic.com/episodes/deepseek-ocr-contexts-optical-compression/
Interactive Visualization: X-LLM: Treating Multimodalities as Foreign Languages