Pete Warden has spent his career on the frontier of small, local AI, first as one of deep learning's earliest engineers (he coined the term “TinyML”) and now as founder of Useful Sensors and Moonshine AI, where he builds voice models that run entirely on-device. Pete joined Ben to make the case that local AI no longer has to be a compromise. They get into what it actually takes to run a capable model on a laptop today; why the voice interface’s bad reputation is a consequence of rough, early implementations rather than a reflection of current capabilities; and where he stands in the ongoing debate between general “end-to-end” models and the compound AI approach of chaining specialized models together. Pete also explains why he thinks browser-based inference could be an "iPhone moment" for local AI and why more and more enterprises are considering self-hosted local models over commercial options. "The shape of [LLMs] is perfect for running locally," Pete says, and local models could be a boon to enterprises worried about cost, privacy, and stability.