This special @Scale Podcast episode is recorded inside Meta’s mechanical and thermal hardware lab, surrounded by next‑generation GPU racks and liquid‑cooling systems that prototype hardware up to five years ahead of deployment.
Joshua Held (Director, Thermal & Mechanical Platform Engineering at Meta) and Yashar Bayani (Director, Production Systems Engineering, Hardware Design at Meta) walk through the progression from simple “pizza box” compute and storage servers to today’s complex GPU racks, explaining how power, cooling, copper limits, and signal integrity now define the cutting edge of AI infrastructure.
They dive into Meta’s journey from buying off‑the‑shelf servers to designing everything end‑to‑end, including the first “Freedom” server, custom data centers like Prineville, and OCP-standard ORV3 racks.
The conversation covers scaling from 8 to 72+ GPUs per rack, the shift to air‑assisted liquid cooling (ALC) for platforms like GB200/GB300, leak detection and resilience, and the manufacturing challenges of thousands of fine connectors, miles of copper, and dust‑sensitive, multimillion‑dollar racks.
Joshua and Yashar also reflect on careers that unexpectedly became “front line” in the AI boom, emphasizing that power is now the primary constraint in hyperscale data centers and that CPUs, GPUs, memory, storage, and optics must be designed holistically. They share how AI agents already help them track work, write docs, and model system designs, and close with advice to students: focus on creative, first‑principles problem solving, grit, and broad foundational skills that stay useful even as AI reshapes entry‑level roles.
Want to join sessions like this in person? Register for our in-person and virtual events on our website: atscaleconference.com