Vanishing Gradients

Episode 54: Scaling AI: From Colab to Clusters — A Practitioner’s Guide to Distributed Training and Inference


Listen Later

Colab is cozy. But production won’t fit on a single GPU.

Zach Mueller leads Accelerate at Hugging Face and spends his days helping people go from solo scripts to scalable systems. In this episode, he joins me to demystify distributed training and inference — not just for research labs, but for any ML engineer trying to ship real software.

We talk through:

• From Colab to clusters: why scaling isn’t just about training massive models, but serving agents, handling load, and speeding up iteration
• Zero-to-two GPUs: how to get started without Kubernetes, Slurm, or a PhD in networking
• Scaling tradeoffs: when to care about interconnects, which infra bottlenecks actually matter, and how to avoid chasing performance ghosts
• The GPU middle class: strategies for training and serving on a shoestring, with just a few cards or modest credits
• Local experiments, global impact: why learning distributed systems—even just a little—can set you apart as an engineer

If you’ve ever stared at a Hugging Face training script and wondered how to run it on something more than your laptop: this one’s for you.

LINKS

  • Zach on LinkedIn
  • Hugo's blog post on Stop Buliding AI Agents
  • Upcoming Events on Luma
  • Hugo's recent newsletter about upcoming events and more!
  • 🎓 Learn more:

    • Hugo's course: Building LLM Applications for Data Scientists and Software Engineershttps://maven.com/s/course/d56067f338
    • Zach's course (45% off for VG listeners!): Scratch to Scale: Large-Scale Training in the Modern World -- https://maven.com/walk-with-code/scratch-to-scale?promoCode=hugo39
    • 📺 Watch the video version on YouTube: YouTube link

      ...more
      View all episodesView all episodes
      Download on the App Store

      Vanishing GradientsBy Hugo Bowne-Anderson

      • 5
      • 5
      • 5
      • 5
      • 5

      5

      11 ratings


      More shows like Vanishing Gradients

      View all
      Data Skeptic by Kyle Polich

      Data Skeptic

      477 Listeners

      a16z Podcast by Andreessen Horowitz

      a16z Podcast

      1,083 Listeners

      The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

      The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

      434 Listeners

      Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

      Super Data Science: ML & AI Podcast with Jon Krohn

      301 Listeners

      NVIDIA AI Podcast by NVIDIA

      NVIDIA AI Podcast

      342 Listeners

      DataFramed by DataCamp

      DataFramed

      268 Listeners

      Practical AI by Practical AI LLC

      Practical AI

      211 Listeners

      Google DeepMind: The Podcast by Hannah Fry

      Google DeepMind: The Podcast

      194 Listeners

      Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

      Machine Learning Street Talk (MLST)

      89 Listeners

      Dwarkesh Podcast by Dwarkesh Patel

      Dwarkesh Podcast

      489 Listeners

      No Priors: Artificial Intelligence | Technology | Startups by Conviction

      No Priors: Artificial Intelligence | Technology | Startups

      131 Listeners

      Latent Space: The AI Engineer Podcast by swyx + Alessio

      Latent Space: The AI Engineer Podcast

      97 Listeners

      AI + a16z by a16z

      AI + a16z

      33 Listeners

      High Signal: Data Science | Career | AI by Delphina

      High Signal: Data Science | Career | AI

      18 Listeners

      OpenAI Podcast by OpenAI

      OpenAI Podcast

      52 Listeners