Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

By Machine Learning Street Talk (MLST)

Welcome! We engage in fascinating discussions with pre-eminent figures in the AI field. Our flagship show covers current affairs in AI, cognitive science, neuroscience and philosophy of mind with in-d... more

  • 4.6
  • 4.6
  • 4.6
  • 4.6
  • 4.6

4.6

95 ratings


Download on the App Store

Best of Machine Learning Street Talk (MLST)

The most played episodes among Podcast App listeners.

  1. Number 1: AI Is Learning at the Wrong Level of Abstraction — Matthieu Wyart

    This episode is sponsored by Notion. Learn more about Notion's Developer Platform today at https://notion.com/mlstWhy can deep networks discover abstractions that shallow models miss? Statistical physicist Matthieu Wyart joins Tim Scarfe to argue that the answer lies in the hidden hierarchy of data. Language and images are built from parts within parts; depth lets a network recover those coarse-grained variables and escape the curse of dimensionality.The conversation moves from jamming transitions and rough loss surfaces to Chomsky, context-free grammars and machine creativity. Wyart explains why next-token prediction can still recover compositional structure, where current systems fall short of genuine scientific invention, and why predicting latent representations rather than raw tokens could make learning far more sample-efficient.They also examine diffusion models, neural scaling laws and the limits of physics-inspired theory. The final question is on a personal note: if mistakes are the price of leaving the beaten path, how much scientific risk is worth taking?---TIMESTAMPS:00:00:00 Can machines learn abstractions from data?00:02:00 Notion agentic workspace00:02:49 From statistical physics to machine learning00:06:40 What physics can explain about learning00:16:37 From Carnot to Chomsky bulldozer00:21:21 How deep networks recover hidden hierarchies00:32:43 Where machine creativity still falls short00:40:48 How deep nets escape the curse of dimensionality00:52:19 Why predict latents instead of tokens01:02:49 The sample-efficiency case for latent prediction01:08:31 Diffusion, scaling laws and text entropy01:16:40 The scientists we learn from and the mistakes we make---REFERENCES:person:[00:00:43] Noam Chomskyhttps://linguistics.mit.edu/user/chomsky/tool:[00:02:08] Notion Developer Platformhttps://www.notion.com/en-gb/blog/introducing-developer-platformpaper:[00:04:43] Mastering the game of Go with deep neural networks and tree searchhttps://www.nature.com/articles/nature16961[00:05:52] Reconciling modern machine-learning practice and the bias-variance trade-offhttps://arxiv.org/abs/1812.11118[00:25:54] How Deep Neural Networks Learn Compositional Data: The Random Hierarchy Modelhttps://arxiv.org/abs/2307.02129[00:42:12] Efficient Estimation of Word Representations in Vector Spacehttps://arxiv.org/abs/1301.3781[00:52:46] Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecturehttps://arxiv.org/abs/2301.08243[00:52:54] Learn from your own latents and not from tokens: A sample-complexity theoryhttps://arxiv.org/abs/2605.27734[01:08:31] A Phase Transition in Diffusion Models Reveals the Hierarchical Nature of Datahttps://arxiv.org/abs/2402.16991[01:11:39] Scaling Laws for Neural Language Modelshttps://arxiv.org/abs/2001.08361[01:12:17] Deriving Neural Scaling Laws from the statistics of natural languagehttps://arxiv.org/abs/2602.07488[01:13:34] Prediction and Entropy of Printed Englishhttps://ieeexplore.ieee.org/document/6773263---LINKS:Download PDF transcript: https://app.rescript.info/share/f7644cdaa86c5cc1e41e484e290f2bd4

    1h 19min
    Listen Later
  2. Number 2: How Deep Learning Finally Cracked Messy Tables - Frank Hutter

    Frank Hutter, co-founder of Prior Labs, talks about TabPFN, a tabular foundation model that makes predictions in a single forward pass, and the research behind it. TabPFN is pre-trained on synthetic datasets drawn from a prior over structural causal models, rather than on real data. At prediction time it takes the whole training table as context and outputs an approximation of the Bayesian posterior predictive distribution, without per-dataset training or hyperparameter search. Frank explains how this grew out of his earlier work on AutoML and neural architecture search, how the priors are built and revised, and why tabular data was hard for deep learning for so long. The conversation also covers the TabArena benchmark, how the architecture changed from TabPFN v1 to v3, scaling to larger tables, using the model with coding agents, test-time compute, Google's TabFM, causal inference and interventions, and relational data. At the end, a short update Frank recorded after the interview covers the TabPFN-3.5 release. Prior Labs: TabPFN-3.5: https://priorlabs.ai/tabpfn-3-5 https://priorlabs.ai/careers TOC: 00:00 Introduction 00:44 Welcome and Frank's background 02:05 Why tabular data was hard for deep learning 10:17 Pre-training on synthetic data 12:52 The TabArena benchmark 19:28 From AutoML to neural architecture search 26:34 TabPFN as a learned algorithm 30:50 Bayesian prediction in one forward pass 39:37 Scaling to larger tables 47:48 Using TabPFN with coding agents 57:47 Output heads and architecture from v1 to v3 1:05:29 Test-time compute and adaptation 1:13:32 Google's TabFM 1:16:53 How the priors are designed 1:18:40 Correlation, causation and interventions 1:35:22 Relational and multimodal data 1:38:31 Use in organisations 1:46:38 The open research arm 1:50:21 Update: TabPFN-3.5 REFS: TabPFN v2, Nature (Hollmann et al., 2025) https://www.nature.com/articles/s41586-024-08328-6 Transformers Can Do Bayesian Inference (Müller et al.) https://arxiv.org/abs/2112.10510 TabArena (Erickson et al.) https://arxiv.org/abs/2506.16791 AutoGluon-Tabular (Erickson et al.) https://arxiv.org/abs/2003.06505 Beyond IID: How General Are Tabular Foundation Models, Really? https://arxiv.org/abs/2606.30410 Neural Architecture Search: A Survey (Elsken, Metzen & Hutter) https://arxiv.org/abs/1808.05377 Auto-WEKA (Thornton et al.) https://www.cs.ubc.ca/~hutter/papers/AutoWEKA-KDD2013.pdf TabPFN v1 (Hollmann et al., 2022) https://arxiv.org/abs/2207.01848 TabPFN-3 technical report https://arxiv.org/abs/2605.13986 TabPFN-2.5 report https://arxiv.org/abs/2511.08667 CAAFE (Hollmann et al.) https://arxiv.org/abs/2305.03403 TabICL (Qu et al.) https://arxiv.org/abs/2502.05564 TabICLv2 (Qu et al.) https://arxiv.org/abs/2602.11139 Google TabFM https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/ TALENT benchmark (Ye et al.) https://arxiv.org/abs/2407.00956 Do-PFN (Robertson et al.) https://arxiv.org/abs/2506.06039 CausalPFN (Balazadeh et al.) https://arxiv.org/abs/2506.07918 Causal Foundation Models with Partial Graphs (Reuter et al.) https://arxiv.org/abs/2602.14972 RelBench (Robinson et al.) https://arxiv.org/abs/2407.20060 RelArena-α, TabPFN-Rel and RPI https://arxiv.org/abs/2608.16319 TabPFN on GitHub https://github.com/PriorLabs/TabPFN TabPFN-3.5 technical report https://arxiv.org/abs/2609.17895 Otto Group Product Classification Challenge (Kaggle, 2015) https://www.kaggle.com/competitions/otto-group-product-classification-challenge ---RESCRIPT:https://app.rescript.info/share/e99676c25ee6189fbf54c9be07eb623e

    1h 54min
    Listen Later
  3. Number 3: Why Scaling Prediction Cannot Create Intelligence - Alexander Mattick

    Alexander Mattick is a researcher at Fraunhofer IIS and a PhD researcher at the University of Technology Nuremberg (UTN), and a regular on Yannic Kilcher's Discord. He first came on MLST in 2022, after helping research the Yann LeCun and Randall Balestriero episode on interpolation. SPONSOR: --- Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Applications for Batch 1 are now open. Apply now: https://cyber.fund --- Alexander treats inference as the thread running through modern machine learning: once you have a model, what does it cost to get an answer out of it? He works through Monte Carlo, GFlowNets, energy-based models, diffusion, normalising flows and flow matching, with four short explainers he recorded himself. He is blunt about energy-based models: you can sample from them in principle, but it is rarely worth the compute. JEPA and "world model", he says, are closer to branding than to technical categories. Next: theories of deep learning, none of which he thinks predicts enough yet to guide practice, then reinforcement learning. --- 0:00 Cold open: information is expensive 0:51 Welcome back, Alexander Mattic 2:08 Alexander's research background 2:50 Inference: densities, sampling and Monte Carlo 6:42 GFlowNets, energy functions and MCMC 9:45 Explainer: energy-based models 11:03 Why model a density at all? 17:30 From learned energies to flow matching 25:08 Explainers: diffusion and normalising flows 28:33 Are energy-based models generative? 33:22 JEPA, contrastive learning and collapse 41:13 Why non-language modalities need flows 44:51 Inference as search: branch and bound 49:43 Q-learning and delayed consequences 55:14 Flow matching, optimal transport, Fokker-Planck 1:00:03 Explainer: flow matching 1:01:49 AlphaFold, latents and scale versus architecture 1:07:52 Two families of deep learning theory 1:15:04 What a good theory would predict 1:23:53 The manifold hypothesis and compression 1:28:25 Is reward enough? 1:32:01 Control theory versus reinforcement learning 1:37:22 The Bitter Lesson and expensive information 1:42:08 Constrained RL: the constrained MDP toolbox 1:50:12 Creativity as constrained search 1:55:44 Reality is protean: when abstractions hold 2:00:32 What is a world model? 2:04:38 Prediction is not control 2:08:13 Robot demos, MPC and reliability --- REFERENCES: [6:55] GFlowNets (Bengio et al., 2021) https://arxiv.org/abs/2106.04399 [38:46] Contrastive Self-Supervised Learning (Anand, 2020) https://ankeshanand.com/blog/2020/01/26/contrative-self-supervised-learning.html [38:56] LeJEPA (Balestriero and LeCun, 2025) https://arxiv.org/abs/2511.08544v3 [47:10] RL for Node Selection in Branch-and-Bound (Mattick) https://openreview.net/forum?id=0ez68a5UqI [56:20] Flow Matching for Generative Modeling https://arxiv.org/abs/2210.02747v2 [1:12:41] Disentangling feature and lazy training in deep neural networks https://arxiv.org/abs/1906.08034v4 [1:31:05] Reward is enough (Silver) https://doi.org/10.1016/j.artint.2021.103535 [1:35:12] Learning ReLU networks to high uniform accuracy is intractable (Berner et al.) https://arxiv.org/abs/2205.13531v2 [1:40:20] Dota 2 with Large Scale Deep RL https://arxiv.org/abs/1912.06680v1 [1:45:41] Constrained Update Projection for Safe Policy Optimization (Yang et al., 2022) https://arxiv.org/abs/2209.07089 [1:46:11] SafeMPO (ICLR 2026) https://openreview.net/forum?id=1m0EU6QXj6 [1:50:17] Why Creativity Cannot Be Interpolated https://archive.mlst.ai/paper/why-creativity-cannot-be-interpolated/ [1:51:39] Invalid Action Masking (Huang and Ontañón) https://arxiv.org/abs/2006.14171 [2:00:04] Probability Theory: The Logic of Science (Jaynes, 2003) https://www.cambridge.org/core/books/probability-theory/9CA08E224FF30123304E6D8935CF1A99 [2:01:53] Training Agents Inside of Scalable World Models (Hafner et al., 2025) https://arxiv.org/abs/2509.24527v1 [2:03:43] World Models (Ha and Schmidhuber, 2018) https://arxiv.org/abs/1803.10122v4

    2h 15min
    Listen Later
  4. Number 4: How Physical AI Learns Across Language, Video and Action — Ming-Yu Liu

    The car making a left turn at the start of this episode was never filmed. Cosmos 3 generated it. Ming-Yu Liu, who leads the Cosmos research at NVIDIA, explains how one model can describe a video, generate one, and produce robot actions. He walks Tim through the architecture. A vision language model reasons one token at a time; its weights then initialise a bidirectional diffusion generator for video, audio and action, and a shared temporal position scheme lines up signals that run at different rates. Ming-Yu treats "world model" as a set of tools, not one definition: forward dynamics, inverse dynamics and policy, trained together under a capacity limit so that each helps the others. He also explains why plentiful first-person human video carries over to robots, which have far less data of their own, and why a Cosmos model post-trained on the DROID dataset is a good starting point for pick-and-place policies. The most practical thread is testing. A neural simulator does not need accurate success rates. It only needs to rank policy A above policy B the way the real world would, so a team can narrow down which checkpoints deserve a real trial. Cosmos Dreams applies that closed-loop idea to driving and robotics, and Ming-Yu argues that humanoids around children and pets make safety matter even more than it does for cars. The conversation ends on the Super, Nano and Edge sizes (Edge targets Jetson Thor, Orin and DGX Spark) and where to find the open weights, code and data. This episode is a paid partnership with NVIDIA. Learn more about Cosmos: https://nvda.ws/4cJoY1S Explore Cosmos Lab: https://research.nvidia.com/labs/cosmos-lab/cosmos3/ --- TIMESTAMPS: 00:00:00 A road that was never filmed 00:02:28 Inside Cosmos 3: reasoning and generator towers 00:05:02 World models: dynamics, policy and one clock 00:08:59 Learning robot skills from human video 00:11:06 Ambiguous tasks and system 2 planning 00:12:53 Neural simulators for policy verification 00:16:41 Cosmos as a starting point for robot policies 00:19:00 Cosmos Dreams and robot safety 00:22:04 Super, Nano and Edge model sizes 00:24:24 Open models, the Cosmos repo and feedback --- REFERENCES: tool: [00:00:13] Cosmos 3 (NVIDIA Cosmos Lab project page) https://research.nvidia.com/labs/cosmos-lab/cosmos3/ [00:18:27] NVIDIA Cosmos GitHub repository https://github.com/NVIDIA/cosmos [00:22:05] Cosmos3-Edge model card https://huggingface.co/nvidia/Cosmos3-Edge [00:22:15] Cosmos3-Super model card https://huggingface.co/nvidia/Cosmos3-Super [00:22:16] Cosmos3-Nano model card https://huggingface.co/nvidia/Cosmos3-Nano [00:22:50] NVIDIA Jetson Thor https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-thor/ [00:22:52] NVIDIA Jetson Orin https://www.nvidia.com/en-us/autonomous-machines/embedded-systems/jetson-orin/ [00:22:53] NVIDIA DGX Spark https://www.nvidia.com/en-us/products/workstations/dgx-spark/ [00:24:42] Cosmos 3 collection on Hugging Face https://huggingface.co/collections/nvidia/cosmos3 other: [00:01:07] Cosmos-Dreams closed-loop simulators (NVIDIA SIGGRAPH 2026 blog) https://blogs.nvidia.com/blog/siggraph-news-2026/ paper: [00:08:54] Cosmos 3: Omnimodal World Models for Physical AI https://arxiv.org/abs/2606.02800 [00:17:43] DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset https://arxiv.org/abs/2403.12945 --- RESCRIPT: https://app.rescript.info/share/e2385948cf465f0d6a2c0930150fc3ab

    26min
    Listen Later
  5. Number 5: AI 2040: Plan A report - Daniel Kokotajlo & Thomas Larsen

    Could slowing AI development make superintelligence safer? Daniel Kokotajlo and Thomas Larsen of the AI Futures Project join Tim Scarfe to examine AI 2040: Plan A, a proposal to buy time before AI exceeds human control. SPONSOR: --- Cyber Fund built the Monastery to help founders ship products that were impossible a year ago. Apply now: https://cyber.fund --- After revisiting AI 2027 and the limits of forecasting, they ask what happens when AI can automate research and sustain an economy without human workers. Tim challenges the case for general models and asks whether intelligence alone explains power. Plan A proposes an initial pause to build safety infrastructure, then cautious development up to the strongest AI that can still be reliably controlled. The discussion tests the distinction between control and alignment, the case for public AI research, and whether the US and China could enforce a slowdown. It ends with the evidence that would change their forecasts. --- TIMESTAMPS: 00:00:00 AI 2040: a slower route to superintelligence 00:01:34 Sponsor: Cyber Fund 00:02:12 From OpenAI to AI 2027 00:06:58 Forecasts, war games and self-fulfilling prophecies 00:17:44 Why AI sceptics are changing their minds 00:23:04 When AI can replace its own researchers 00:28:45 Could an AI economy grow without human workers? 00:37:32 One general model or a society of specialists? 00:47:43 Brains, machines and collective intelligence 00:56:12 Plan A: buy time at the controllable frontier 01:00:02 Why control buys time but cannot replace alignment 01:06:36 Why AI research should be public 01:10:32 Can the US and China enforce an AI slowdown? 01:19:04 Why AI policy debates miss the technology 01:21:56 Is AI normal technology? The remaining disagreement Many thanks to James Wilken-Smith for helping with show research. --- REFERENCES: other: [00:00:01] AI 2040: Plan A https://ai-2040.com/ [00:03:27] AI 2027 https://ai-2027.com/ [00:13:47] Scenario Scrutiny for AI Policy https://blog.aifutures.org/p/scenario-scrutiny-for-ai-policy [00:33:11] The 2028 Global Intelligence Crisis https://www.citriniresearch.com/p/2028gic [01:00:40] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident https://www.redwoodresearch.org/research/hugging-face-incident [01:09:21] The Hugging Face incident and the road ahead https://openai.com/index/hugging-face-incident-and-the-road-ahead/ [01:22:01] AI as Normal Technology https://www.normaltech.ai/p/ai-as-normal-technology [01:22:51] Common Ground between AI 2027 & AI as Normal Technology https://asteriskmag.substack.com/p/common-ground-between-ai-2027-and person: [00:19:43] Geoffrey Hinton https://www.cs.toronto.edu/~hinton/ [00:20:07] Ryan Greenblatt https://www.lesswrong.com/users/ryan_greenblatt [00:26:06] Elon Musk https://www.tesla.com/elon-musk tool: [00:21:46] ARC-AGI-3 https://arcprize.org/arc-agi/3 [00:21:53] AlphaGo and Move 37 https://deepmind.google/research/alphago/ [00:39:41] Claude https://claude.com/product/overview [00:39:58] NVIDIA H100 GPU https://www.nvidia.com/en-us/data-center/h100/ paper: [00:24:42] Training AI Scientists to Replicate Research https://arxiv.org/abs/2608.13331v1 [01:27:19] Validity of the single processor approach to achieving large scale computing capabilities https://www.cs.cmu.edu/~18742/papers/Amdahl1967.pdf book: [00:28:52] Bullshit Jobs: A Theory https://www.simonandschuster.com/books/Bullshit-Jobs/David-Graeber/9781501143335 organization: [01:05:09] Redwood Research https://www.redwoodresearch.org/ --- RESCRIPT: https://app.rescript.info/public/share/33d1a58fa8f307ae7dfd504d4fdaa9d5

    1h 30min
    Listen Later

Machine Learning Street Talk (MLST) episodes:

FAQs about Machine Learning Street Talk (MLST):

How many episodes does Machine Learning Street Talk (MLST) have?

The podcast currently has 270 episodes available.

More shows like Machine Learning Street Talk (MLST)

The a16z Show by Andreessen Horowitz

The a16z Show

1,092 Listeners

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

439 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

304 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

207 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

205 Listeners

Last Week in AI by Skynet Today

Last Week in AI

317 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

591 Listeners

Big Technology Podcast by Alex Kantrowitz

Big Technology Podcast

503 Listeners

No Priors: Artificial Intelligence | Technology | Startups by Conviction

No Priors: Artificial Intelligence | Technology | Startups

143 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

102 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

685 Listeners

BG2Pod with Brad Gerstner and Bill Gurley by BG2Pod

BG2Pod with Brad Gerstner and Bill Gurley

456 Listeners

AI + a16z by a16z

AI + a16z

30 Listeners