
Sign up to save your podcasts
Or


What does it take to build and scale the infrastructure powering the AI revolution?
In this episode of the @Scale podcast, Francois Richard, Engineering Director at Meta, sits down with Jim Julson, VP of Global Infrastructure Networking at CoreWeave, to explore the technical and organizational challenges of scaling AI infrastructure.
Jim shares lessons from his experience helping CoreWeave grow from standing up four data centers in just 90 days to operating a global infrastructure footprint that continues to evolve alongside the rapidly growing demands of AI.
The conversation covers:
• Building and scaling AI infrastructure at unprecedented speed
• The role of engineering “heroes” — and why teams need a culture of ownership to scale effectively
• Event-driven automation and CoreWeave’s approach to infrastructure management
• Building observability systems for large-scale AI infrastructure
• Capacity planning across GPUs, CPUs, switches, and networks
• The opportunities and limitations of using AI agents in production environments
• Why infrastructure architectures need to continuously evolve as AI workloads grow
• Building an engineering culture where teams collaborate, take ownership, and learn from one another
Jim also shares a look at his personal passion for automation: building a fully automated home theater with custom LED lighting and event-driven controls.
Tune in for a behind-the-scenes look at the infrastructure, engineering practices, and culture helping power AI at scale.
Want to attend conversations like this in person? Register for an @Scale event here: atscaleconference.com
Anoop leads applied science for AWS's Agentic AI unit, building autonomous agents for both business productivity and software development. In this episode, he traces his path from speech recognition research to leading AWS Transform and Kiro, AWS's coding agent.
He digs into how agents learn organizational conventions without being explicitly taught, why memory management is one of the hardest unsolved problems in agentic AI, and why the next frontier is proactive, not just reactive, agents. Also: lessons from Netflix and Microsoft on personalization, and why code review (not writing code) is now the real bottleneck.
Anoop is joined by @Scale Podcast host, Francois Richard — Engineering Director responsible for the Reliability Infra at Meta.
Want to join sessions like this in person?
Register for our in-person and virtual events on our website: atscaleconference.com
From air-cooled gear to liquid-cooled GPU racks pulling hundreds of kilowatts, Meta’s AI infrastructure is transforming what it means to be a thermal mechanical engineer.
As Joshua Held shares, the AI renaissance pulled his team into the front lines of designing rack-scale systems, liquid cooling, and dense GPU deployments, turning what once felt like a niche specialty into a creative awakening.
In this @Scale episode with Joshua Held and Yashar Bayani, hosted by Francois Richard, they dive into GPU racks, liquid cooling, power constraints, and how Meta is pushing copper and optics to their limits to build next-generation rack-scale AI systems.
This special @Scale Podcast episode is recorded inside Meta’s mechanical and thermal hardware lab, surrounded by next‑generation GPU racks and liquid‑cooling systems that prototype hardware up to five years ahead of deployment.
Joshua Held (Director, Thermal & Mechanical Platform Engineering at Meta) and Yashar Bayani (Director, Production Systems Engineering, Hardware Design at Meta) walk through the progression from simple “pizza box” compute and storage servers to today’s complex GPU racks, explaining how power, cooling, copper limits, and signal integrity now define the cutting edge of AI infrastructure.
They dive into Meta’s journey from buying off‑the‑shelf servers to designing everything end‑to‑end, including the first “Freedom” server, custom data centers like Prineville, and OCP-standard ORV3 racks.
The conversation covers scaling from 8 to 72+ GPUs per rack, the shift to air‑assisted liquid cooling (ALC) for platforms like GB200/GB300, leak detection and resilience, and the manufacturing challenges of thousands of fine connectors, miles of copper, and dust‑sensitive, multimillion‑dollar racks.
Joshua and Yashar also reflect on careers that unexpectedly became “front line” in the AI boom, emphasizing that power is now the primary constraint in hyperscale data centers and that CPUs, GPUs, memory, storage, and optics must be designed holistically. They share how AI agents already help them track work, write docs, and model system designs, and close with advice to students: focus on creative, first‑principles problem solving, grit, and broad foundational skills that stay useful even as AI reshapes entry‑level roles.
Want to join sessions like this in person? Register for our in-person and virtual events on our website: atscaleconference.com
In Episode 5 of the @Scale Podcast, Danran Chen, Senior Product Manager of AI at Zoom, shares how artificial intelligence is transforming product development and collaboration.
Danran discusses her path from engineering at Airbnb to leading AI products at Zoom, how AI tools are reshaping the role of product managers, and why building systems around large language models is key to delivering real user value.
She also explains Zoom’s federated AI strategy and how AI can free people from repetitive work, allowing more time for creativity and meaningful human connection.
Danran is joined by our host, Francois Richard, Engineering Director at Meta.
Learn more about the @Scale Conference here: https://bit.ly/4sIjGcB
In this podcast episode, Jessica Powell, co-founder and CEO of Audio Shake, discusses her company's innovative AI audio separation technology and journey from Google to startup founder.
Jessica Powell is the CEO and co-founder of AudioShake, a sound-splitting AI technology that makes audio more usable for both humans and machines.
Francois Richard is Engineering Director responsible for the Reliability Infra at Meta. Reliability Infra is focused on improving Meta’s reliability across the entire lifecycle of incidents by ensuring that Meta can swiftly and confidently recover from any type of outage caused by both known and unknown failures. Francois started at Facebook in 2017. His career spans nearly two decades of working in speech recognition, search, e-mail systems and distributed systems at Nuance, Yahoo and Meta. He holds degrees in Electrical Engineering and Computer Science from Ecole Polytechnique of Montreal.
🎙️ New Episode: @Scale Podcast, Ep. 3
This week, Dr. Pradeep Sindhu — industry visionary, founder of Juniper Networks, inventor of the DPU, and current Microsoft leader — joins host Francois Richard of Meta for an insightful conversation.
You’ll hear about:
✨ Lessons from Xerox PARC, Sun Microsystems, and Juniper
✨ The invention of the DPU and why CPUs can’t handle modern workloads
✨ How today’s AI moment compares to the rise of the internet
✨ Advice for startups, scaling teams, and innovating in hardware
Perfect for senior engineers, startup builders, and anyone curious about the future of computing and AI infrastructure.
Dr. Pradeep Sindhu is an industry visionary currently focused on data processing innovations at Microsoft. Sindhu, who co-founded Fungible Inc. and served as its CEO and CTO, is credited with inventing the Data Processing Unit (DPU) that revolutionized storage system efficiency. He is also the founder of Juniper Networks, where he led the development of all major products that shaped the future of networking infrastructure. Dr. Sindhu’s contributions have redefined networking hardware and software, driving advances impacting cloud computing and AI infrastructure.
Francois Richard is Engineering Director responsible for the Reliability Infra at Meta. Reliability Infra is focused on improving Meta’s reliability across the entire lifecycle of incidents by ensuring that Meta can swiftly and confidently recover from any type of outage caused by both known and unknown failures. Francois started at Facebook in 2017. His career spans nearly two decades of working in speech recognition, search, e-mail systems and distributed systems at Nuance, Yahoo and Meta. He holds degrees in Electrical Engineering and Computer Science from Ecole Polytechnique of Montreal.
Joe Spisak, Product Director for Artificial Intelligence at Meta, sits down with Ion Stoica to discuss a wide range of topics related to the AI and machine learning industry, including the resurgence of reinforcement learning, large language models, the power of open source software, and the evolution of the AI tech stack.
Ion Stoica is a Professor in the EECS Department at the University of California at Berkeley, and the Director of SkyLab (https://sky.cs.berkeley.edu/). He is currently doing research on cloud computing and AI systems. Past work includes Ray, Apache Spark, Apache Mesos, Tachyon, Chord DHT, and Dynamic Packet State (DPS). He is an Honorary Member of the Romanian Academy, an ACM Fellow and has received numerous awards, including the Mark Weiser Award (2019), SIGOPS Hall of Fame Award (2015), and several "Test of Time" awards. He also co-founded three companies, Anyscale (2019), Databricks (2013) and Conviva (2006).Joe Spisak is Product Director and Head of Open Source in Meta’s Generative AI organization. A veteran of the AI space with over 10 years experience, Joe led product teams at Meta/Facebook, Google and Amazon where he focused on open source AI, open science and building developer tools such as PyTorch to help the community scale up AI in an open and collaborative way. As the leader of product for PyTorch, he and the team built an amazing platform and community that made PyTorch the leading open source AI development framework in the industry and took it to the Linux Foundation where it now resides in partnership with Microsoft, Nvidia, AMD, Google, Amazon and more. Joe is also an angel and advisor to companies like Anthropic, Answer.ai, Lastmile.ai, Evolutionary Scale, Udacity, Lightning.AI, and others.
Learn more about @Scale here: https://atscaleconference.com/
In this podcast, Mohamed Fawzy, Senior Director of Engineering at NVIDIA, joins our host Francois Richard to discuss stories and insights from running hyper-scale systems and AI innovation at companies like Meta, Cruise, Yahoo and NVIDIA.
Francois Richard is Engineering Director responsible for the Reliability Infra at Meta. Reliability Infra is focused on improving Meta’s reliability across the entire lifecycle of incidents by ensuring that Meta can swiftly and confidently recover from any type of outage caused by both known and unknown failures. Francois started at Facebook in 2017. His career spans nearly two decades of working in speech recognition, search, e-mail systems and distributed systems at Nuance, Yahoo and Meta. He holds degrees in Electrical Engineering and Computer Science from Ecole Polytechnique of Montreal.
Mohamed Fawzy Senior Director of Engineering at NVIDIA, works on building large-scale AI infrastructure to enhance researcher productivity and optimize GPU computing at NVIDIA. With a background in distributed systems and high-performance computing, he focuses on developing scalable, efficient, and reliable infrastructure that accelerates AI workloads. His work involves improving system performance, streamlining AI workflows, and enabling researchers to iterate faster on complex models.
From the publisher's feed