Super Data Science: ML & AI Podcast with Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

By Jon KrohnScienceTechnology
Download on the App Store
  • Favorites

    271

    Followers

  • Typical duration

    55 min

    per episode

Based on Podcast App listening data

Super Data Science: ML & AI Podcast with Jon Krohn episodes

  • 1032: Agents Need 10x More Data Than Humans, with Salesforce’s CDO Michael Andrew

    During their #sponsored discussion, Chief Data Officer at Salesforce Michael Andrew talks to Jon Krohn about what changes for a data team when its customers are AI agents as well as people. Listen to the episode to hear Michael Andrew talk about why agents need ten times more trusted data than humans, how Salesforce untrapped its own customer data with Data 360 and what practitioners should be learning to stay effective in the agentic era!


    Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1032⁠⁠


    Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.⁠⁠⁠


    In this episode you will learn:

    • (02:32) How the CDO role changes when agents are customers
    • (06:18) Why agents need ten times more data than humans
    • (09:06) How Salesforce untrapped its own customer data
    • (21:36) What practitioners should be learning for the agentic era
    • 28 min
    • 1031: Tokenomics: Why Your Agentic AI Bill Is Exploding (and How to Fix It), with Tyler Cox and Ish Shah

      In Episode #1031, Ish Shah and Tyler Cox (Distinguished Engineers in the Office of the CTO for Dell Technologies' client group) join Jon Krohn to work out why agentic AI bills are exploding and what can be done about it. Over one weekend Ish burned roughly two billion tokens on a side project, and that is the ordinary shape of agentic work now: agents spawn sub-agents, the pie of work grows, and cheaper tokens only invite more ambitious projects. Tyler runs a small Dell lab that pushes hundreds of millions of tokens a day through local hardware instead.


      In this episode, they define what makes a system agentic, explain how to read a Pareto curve when choosing models, work through the jagged frontier and why most tasks do not need a frontier model, and lay out what moving agentic workloads onto your own hardware does to the economics.


      Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1031⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


      Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


      In this episode you will learn:

      • (00:03:42) What makes a system agentic
      • (00:12:45) Picking the right model for the task
      • (00:16:46) How to read a Pareto curve
      • (00:27:02) Why agents burn so many more tokens
      • 1 hr 16 min
      • 1030: Garbage In, Gospel Out: Why Agents Need Better Data, with Salesforce's Gaurav Pathak

        During their #sponsored discussion, Senior Vice President Product Management AI and Metadata at Salesforce, Gaurav Pathak talks to Jon Krohn about why AI agents need well-labeled, high-quality data to deliver reliable answers in the enterprise. Listen to the episode to hear Gaurav Pathak talk about the difference between a “data brawl” and “garbage in, gospel out”, who the “sin eaters” of enterprise AI are and the three skills that matter most for AI engineers today!


        Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1030⁠


        Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.⁠⁠⁠


        In this episode you will learn:

        • (03:05) Why metadata are the labels AI agents need
        • (06:31) From “data brawl” to “garbage in, gospel out”
        • (10:57) Who the “sin eaters” of enterprise AI are
        • (13:45) What data quality rules are and how CLAIRE generates them
        • (17:21) Three skills AI engineers need in the agentic era
        • 23 min
        • 1029: How AI Brought a Podcast Back From the Dead, with Linear Digressions’ Katie Malone

          In Episode #1029, Dr. Katie Malone (Host of Linear Digressions) joins Jon Krohn to explain how AI brought her podcast back from the dead. After nearly 300 episodes, Katie shut down Linear Digressions due to burnout, but better tools helped her relaunch it six years later. Along the way she has taught machine learning at Udacity and the University of Chicago and led the development of agentic AI platforms inside a company of tens of thousands of people. In this episode, she argues that people management and agent management are the same skill in different clothing, works through what AI slop and process slop are doing to organisations, describes the agent that now produces her show, and takes a pop quiz on three of her favourite data paradoxes.


          Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1027⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


          Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


          In this episode you will learn:

          • (00:04:01) Why Linear Digressions stopped, and what changed enough to bring it back
          • (00:15:01) Why people management and agent management are the same skill
          • (00:24:32) The "Claude Code in a trench coat" agent that produces her show
          • (00:41:39) Bainbridge’s ironies of automation, and why expertise gets rusty
          • 1 hr 11 min
          • 1028: The Chip Built for Agentic AI Inference, with SambaNova's Anton McGonnell

            In Episode #1028, Anton McGonnell (VP of Product at SambaNova) joins Jon Krohn to explain why the chips running most AI inference today were never designed for the job. Agentic AI has changed the computational profile of inference, with much larger inputs and far heavier caches feeding the token generation that follows, and that shift has exposed where GPU architecture struggles. SambaNova has raised over $2 billion to build an alternative, the reconfigurable dataflow unit, which lays a whole model out spatially across the chip rather than executing it kernel by kernel. In this episode, Anton discusses why the speed that matters is payback, and how speed and concurrency are what turn a fixed hardware cost into a six-month payback. He also walks through the trade-off every inference provider faces between speed per user and throughput per chip, what the RDU architecture changes about scaling and data center deployment, the economics of the new SN50, and why four out of five AI infrastructure leaders say they would pay a premium for faster tokens.


            Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1028⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


            Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


            In this episode you will learn:

            • (00:02:41) Why agentic AI is reshaping inference workloads
            • (00:08:39) How SambaNova's RDU differs from a GPU
            • (00:17:54) The economics of the SN50
            • 30 min
            • 1027: Building an Always-On AI Agent for Busy Parents, with Dr. Dilani Kahawala

              In Episode #1027, Dr. Dilani Kahawala (Co-Founder and CEO of Anna) joins Jon Krohn to explain what it takes to build an always-on AI assistant that busy parents will trust with their inboxes. Anna watches the email, school apps, WhatsApp messages and calendars flowing into a family's life and surfaces what matters, over text and voice, with barely any app to speak of. Dilani came to it by way of a Harvard physics PhD, McKinsey, and a decade of product leadership at Etsy, Meta and Atlassian, and says she has had to throw away most of what that decade taught her about how products get built. In this episode, she lays out the three hardest problems in building Anna, why the eval loop is the heart of the product, how a long-running agent differs from a turn-based one, and the brutal unit economics of consumer AI.


              Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1027⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


              Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


              In this episode you will learn:

              • (00:10:01) The three hardest problems in building a consumer agent
              • (00:13:23) Why a long-running agent is a different problem from a turn-based one
              • (00:22:33) Why the eval and improvement loop is the heart of the product
              • (00:26:42) The unit economics of always-on AI on a flat subscription
              • 1 hr 3 min
              • 1026: OpenAI’s GPT-6 Astra

                In Episode #1026, Jon Krohn breaks down GPT-6 Astra, OpenAI’s new flagship that its president has floated as a possible marker of AGI. Jon covers what the model is, what it costs, its state-of-the-art results across computer use, coding, abstract reasoning and science and the safety story, which for this release is unusually intertwined with capability. He weighs the AGI claim against Anthropic’s Fable 5.1 and lands, as ever, in a measured middle.


                Additional materials:⁠ ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1026⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


                Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


                In this episode you will learn:

                • (00:14) What GPT-6 Astra is, and what it costs

                • (03:58) The capability highlights that matter most

                • (10:23) The safety story and the AGI question

                • 20 min
                • 1025: Word Gravity: How Transformers Bend Space, with Dr. Luis Serrano

                  In Episode #1025, Dr. Luis Serrano (Founder of Serrano Academy) joins Jon Krohn to explain the paper he co-authored on the curved spacetime of transformer architectures, in which attention stops being a lookup table and becomes something closer to gravity: words bend the space around them, and the embedding of "bank" visibly curves toward "river" as it travels through the layers of the network. In this episode, he recreates Eddington’s 1919 eclipse experiment inside a transformer, draws the line between an LLM workflow and an actual agent, explains why agent evaluation is a step harder than evaluating an essay, and gives the cleanest account of GRPO you will hear.


                  Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1025⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


                  Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


                  In this episode you will learn:

                  • (00:10:53) What changed, and what survived, between the two editions of Grokking Machine Learning
                  • (00:27:48) Word gravity: how attention pulls "bank" toward "river"
                  • (00:42:12) Why RAG is an LLM workflow rather than an agent
                  • (00:51:23) The two-by-two that explains why GRPO powers reasoning models
                  • 1 hr 11 min
                  • 1024: In Case You Missed It in August 2026

                    In ICYMI Episode #1024, Jon Krohn tracks the gap between AI investment and AI return, from the technology side to the people side. Hear from Pete Johnson, Jerry Yurchisin, Priyanka Vergadia and Tristan Handy, discussing why four out of five organizations have the structures for AI success in place while only one in five sees the returns, which decisions should never be handed to a language model however confident it sounds, how to structure Claude skills so that your output stops being slop and why the semantic layer matters more, not less, now that analytics agents are the ones asking the questions.


                    Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠www.superdatascience.com/1024⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


                    Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


                    In this episode you will learn:

                    • (00:56) Vector Search, Agentic Memory and Effective RAG
                      • (09:20) Mathematical Optimization in the Agentic AI Era
                        • (17:30) Anyone Can Write Code Now, So What Gets You Hired?
                          • (27:14) How dbt Won Analytics Engineering
                          • 35 min
                          • 1023: Agentic AI Skills That Matter Now, with Aishwarya Srinivasan

                            In Episode #1023, Aishwarya Srinivasan (Co-Founder of The Gen Academy) joins Jon Krohn to work out where a competitive moat comes from once anything you can build in ten minutes, somebody else can build in ten minutes too. Ash came to teaching through Illuminate AI, the mentorship community she started in 2020, and now trains senior engineers and leaders to ship agentic AI in production; she is blunt that vibe coding lowers the floor without touching the engineering judgment that production demands. In this episode, she explains what a whole-system eval covers that a model eval misses, traces reinforcement learning from the algorithm she patented at IBM to its resurgence in agentic fine tuning and lays out the MIND framework from her TED Talk for living with AI.


                            Additional materials: ⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠https://www.superdatascience.com/1023⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠⁠


                            Interested in sponsoring a SuperDataScience Podcast episode? Email [email protected] for sponsorship information.


                            In this episode you will learn:

                            • (00:10:10) Why cheap code shifts the software engineering job rather than ending it
                            • (00:15:50) What a whole-system eval covers that a model eval misses
                            • (00:36:11) Why reinforcement learning came roaring back for agentic AI
                            • (00:41:23) The one skill Ash says matters more than any hard skill
                            • 1 hr 19 min

                            About Super Data Science: ML & AI Podcast with Jon Krohn

                            From the publisher's feed

                            The latest machine learning, A.I., and data career topics from across both academia and industry are brought to you by host Dr. Jon Krohn on the Super Data Science Podcast. As the quantity of data on our planet doubles every couple of years and with this trend set to continue for decades to come, there's an unprecedented opportunity for you to make a meaningful impact in your lifetime. In conversation with the biggest names in the data science industry, Jon cuts through hype to fuel that professional impact.

                            Best of Super Data Science: ML & AI Podcast with Jon Krohn

                            Ranked by our users in the last 21 days

                            More shows like Super Data Science: ML & AI Podcast with Jon Krohn

                            Data Skeptic by Kyle Polich

                            Data Skeptic

                            476 Listeners

                            Software Engineering Daily by Software Engineering Daily

                            Software Engineering Daily

                            624 Listeners

                            Talk Python To Me by Michael Kennedy

                            Talk Python To Me

                            582 Listeners

                            NVIDIA AI Podcast by NVIDIA

                            NVIDIA AI Podcast

                            338 Listeners

                            AI Today Podcast by AI & Data Today

                            AI Today Podcast

                            155 Listeners

                            DataFramed by DataCamp

                            DataFramed

                            265 Listeners

                            Practical AI by Daniel Whitenack and Chris Benson

                            Practical AI

                            202 Listeners

                            The Real Python Podcast by Real Python

                            The Real Python Podcast

                            139 Listeners

                            Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

                            Machine Learning Street Talk (MLST)

                            99 Listeners

                            No Priors: Artificial Intelligence | Technology | Startups by Conviction

                            No Priors: Artificial Intelligence | Technology | Startups

                            140 Listeners

                            AI Chat: AI News & Artificial Intelligence by Jaeden Schafer

                            AI Chat: AI News & Artificial Intelligence

                            164 Listeners

                            This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                            This Day in AI Podcast

                            222 Listeners

                            The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                            The AI Daily Brief: Artificial Intelligence News and Analysis

                            682 Listeners

                            AI For Humans: Weekly AI News, Tools & Trends by Kevin Pereira & Gavin Purcell

                            AI For Humans: Weekly AI News, Tools & Trends

                            273 Listeners

                            Training Data by Sequoia Capital

                            Training Data

                            39 Listeners