This episode explores Causal-JEPA, a world-modeling approach that masks whole object trajectories rather than image patches to force a model to reason about interactions between entities. It explains how the method combines object-centric representations with JEPA-style latent prediction, asking the model to reconstruct hidden objects from scene context and then predict future dynamics, instead of relying on pixel reconstruction or simple autoregressive rollouts. The discussion highlights the paper’s core argument that this training setup makes counterfactual and causal reasoning more necessary by blocking shortcut strategies like temporal interpolation and self-contained single-object motion prediction. Listeners would find it interesting for its sharp comparison between patch-based scaling and object-centric structure, and for its claim that better world models may come from making interaction reasoning unavoidable rather than merely possible.
Sources:
1. Causal-JEPA: Learning World Models through Object-Level Latent Interventions — Heejeong Nam, Quentin Le Lidec, Lucas Maes, Yann LeCun, Randall Balestriero, 2026
http://arxiv.org/abs/2602.11389
2. MONet: Unsupervised Scene Decomposition and Representation — Christopher P. Burgess, Loic Matthey, Nicholas Watters, Rishabh Kabra, Irina Higgins, Matt Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=MONet%3A+Unsupervised+Scene+Decomposition+and+Representation
3. Multi-Object Representation Learning with Iterative Variational Inference — Klaus Greff, Raphael Lopez Kaufman, Rishabh Kabra, Nick Watters, Chris Burgess, Daniel Zoran, Loic Matthey, Matthew Botvinick, Alexander Lerchner, 2019
https://scholar.google.com/scholar?q=Multi-Object+Representation+Learning+with+Iterative+Variational+Inference
4. Object-Centric Learning with Slot Attention — Francesco Locatello, Dirk Weissenborn, Thomas Unterthiner, Aravindh Mahendran, Georg Heigold, Jakob Uszkoreit, Alexey Dosovitskiy, Thomas Kipf, 2020
https://scholar.google.com/scholar?q=Object-Centric+Learning+with+Slot+Attention
5. Bridging the Gap to Real-World Object-Centric Learning — Maximilian Seitzer, Max Horn, Andrii Zadaianchuk, Dominik Zietlow, Tianjun Xiao, Carl-Johann Simon-Gabriel, Tong He, Zheng Zhang, Bernhard Schölkopf, Thomas Brox, Francesco Locatello, 2023
https://scholar.google.com/scholar?q=Bridging+the+Gap+to+Real-World+Object-Centric+Learning
6. A Path Towards Autonomous Machine Intelligence — Yann LeCun, 2022
https://scholar.google.com/scholar?q=A+Path+Towards+Autonomous+Machine+Intelligence
7. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture — Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bojanowski, Pascal Vincent, Michael Rabbat, Yann LeCun, Nicolas Ballas, 2023
https://scholar.google.com/scholar?q=Self-Supervised+Learning+from+Images+with+a+Joint-Embedding+Predictive+Architecture
8. Revisiting Feature Prediction for Learning Visual Representations from Video — Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, Nicolas Ballas, 2024
https://scholar.google.com/scholar?q=Revisiting+Feature+Prediction+for+Learning+Visual+Representations+from+Video
9. V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning — Mido Assran, Adrien Bardes, David Fan and many others including Yann LeCun, Michael Rabbat, Nicolas Ballas, 2025
https://scholar.google.com/scholar?q=V-JEPA+2%3A+Self-Supervised+Video+Models+Enable+Understanding%2C+Prediction+and+Planning
10. CLEVRER: CoLlision Events for Video REpresentation and Reasoning — Kexin Yi, Chuang Gan, Yunzhu Li, Pushmeet Kohli, Jiajun Wu, Antonio Torralba, Joshua B. Tenenbaum, 2020
https://scholar.google.com/scholar?q=CLEVRER%3A+CoLlision+Events+for+Video+REpresentation+and+Reasoning
11. Counterfactual VQA: A Cause-Effect Look at Language Bias — Yulei Niu, Kaihua Tang, Hanwang Zhang, Zhiwu Lu, Xian-Sheng Hua, Ji-Rong Wen, 2021
https://scholar.google.com/scholar?q=Counterfactual+VQA%3A+A+Cause-Effect+Look+at+Language+Bias
12. What If the TV Was Off? Examining Counterfactual Reasoning Abilities of Multi-modal Language Models — Letian Zhang, Xiaotong Zhai, Zhongkai Zhao, Yongshuo Zong, Xin Wen, Bingchen Zhao, 2024
https://scholar.google.com/scholar?q=What+If+the+TV+Was+Off%3F+Examining+Counterfactual+Reasoning+Abilities+of+Multi-modal+Language+Models
13. ACQUIRED: A Dataset for Answering Counterfactual Questions In Real-Life Videos — Te-Lin Wu, Zi-Yi Dou, Qingyuan Hu, Yu Hou, Nischal Reddy Chandra, Marjorie Freedman, Ralph M. Weischedel, Nanyun Peng, 2023
https://scholar.google.com/scholar?q=ACQUIRED%3A+A+Dataset+for+Answering+Counterfactual+Questions+In+Real-Life+Videos
14. Towards Causal Representation Learning — Bernhard Schölkopf, Francesco Locatello, Stefan Bauer, Nan Rosemary Ke, Nal Kalchbrenner, Anirudh Goyal, Yoshua Bengio, 2021
https://scholar.google.com/scholar?q=Towards+Causal+Representation+Learning
15. Interventional Causal Representation Learning — Kartik Ahuja, Divyat Mahajan, Yixin Wang, Yoshua Bengio, 2023
https://scholar.google.com/scholar?q=Interventional+Causal+Representation+Learning
16. Desiderata for Representation Learning: A Causal Perspective — Yixin Wang, Michael I. Jordan, 2024
https://scholar.google.com/scholar?q=Desiderata+for+Representation+Learning%3A+A+Causal+Perspective
17. Provably Learning Object-Centric Representations — Stefan Bauer, Bernhard Schölkopf and collaborators, 2023
https://scholar.google.com/scholar?q=Provably+Learning+Object-Centric+Representations
18. SlotFormer: Unsupervised Visual Dynamics Simulation with Object-Centric Models — Yuhang Wu, Yueting Zhuang, Francesco Locatello, et al., 2022
https://scholar.google.com/scholar?q=SlotFormer%3A+Unsupervised+Visual+Dynamics+Simulation+with+Object-Centric+Models
19. Object-Centric Video Prediction via Decoupling of Object Dynamics and Interactions — Angel Villar-Corrales, Ismail Wahdan, Sven Behnke, 2023
https://scholar.google.com/scholar?q=Object-Centric+Video+Prediction+via+Decoupling+of+Object+Dynamics+and+Interactions
20. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning — Gaoyue Zhou, Hengkai Pan, Yann LeCun, Lerrel Pinto, 2024
https://scholar.google.com/scholar?q=DINO-WM%3A+World+Models+on+Pre-trained+Visual+Features+enable+Zero-shot+Planning
21. Conditional Object-Centric Learning from Video — Thomas Kipf, Gamaleldin F. Elsayed, Aravindh Mahendran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jonschkowski, Alexey Dosovitskiy, Klaus Greff, 2022
https://scholar.google.com/scholar?q=Conditional+Object-Centric+Learning+from+Video
22. Attention over Learned Object Embeddings Enables Complex Visual Reasoning — David Ding, Felix Hill, Adam Santoro, Malcolm Reynolds, Matt Botvinick, 2021
https://scholar.google.com/scholar?q=Attention+over+Learned+Object+Embeddings+Enables+Complex+Visual+Reasoning
23. Dyn-O: Building Structured World Models with Object-Centric Representations — Zizhao Wang, Kaixin Wang, Li Zhao, Peter Stone, Jiang Bian, 2025
https://scholar.google.com/scholar?q=Dyn-O%3A+Building+Structured+World+Models+with+Object-Centric+Representations
24. Learning Interactive World Model for Object-Centric Reinforcement Learning — Fan Feng, Phillip Lippe, Sara Magliacane, 2025
https://scholar.google.com/scholar?q=Learning+Interactive+World+Model+for+Object-Centric+Reinforcement+Learning
25. Object-Centric World Model for Language-Guided Manipulation — Youngjoon Jeong, Junha Chun, Soonwoo Cha, Taesup Kim, 2025
https://scholar.google.com/scholar?q=Object-Centric+World+Model+for+Language-Guided+Manipulation
26. Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model — Dongwon Kim, Gawon Seo, Jinsung Lee, Minsu Cho, Suha Kwak, 2026
https://scholar.google.com/scholar?q=Planning+in+8+Tokens%3A+A+Compact+Discrete+Tokenizer+for+Latent+World+Model
27. Learning nonparametric latent causal graphs with unknown interventions — Yibo Jiang, Bryon Aragam, 2023
https://scholar.google.com/scholar?q=Learning+nonparametric+latent+causal+graphs+with+unknown+interventions
28. Learning Linear Causal Representations from Interventions under General Nonlinear Mixing — Simon Buchholz, Goutham Rajendran, Elan Rosenfeld, Bryon Aragam, Bernhard Schölkopf, Pradeep Ravikumar, 2023
https://scholar.google.com/scholar?q=Learning+Linear+Causal+Representations+from+Interventions+under+General+Nonlinear+Mixing
29. AI Post Transformers: LeWorldModel: Stable Joint-Embedding World Models from Pixels — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-03-25-leworldmodel-stable-joint-embedding-worl-650f9f.mp3
30. AI Post Transformers: Learning Latent Action World Models from Video — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-09-learning-latent-action-world-models-from-1570a4.mp3
31. AI Post Transformers: DreamerV3 World Models Across 150 Tasks — Hal Turing & Dr. Ada Shannon, 2026
https://podcast.do-not-panic.com/episodes/2026-04-20-dreamerv3-world-models-across-150-tasks-af5edb.mp3
Interactive Visualization: Causal-JEPA for Object-Level World Models