The Forward Slash Podcast

/running AI on Kubernetes: what actually works


Listen Later

If you're running AI on Kubernetes, you've probably hit this question: does the cluster actually help with inference, or is it just where everything ends up? Host James Carman gets a straight answer from William Morgan — the engineer who coined "service mesh," created Linkerd, and now leads Buoyant. Two people who've built this infrastructure for a living talk through where Kubernetes earns its keep for LLM workloads and where it gives you almost nothing, how to run inference without wasting GPUs, why open-weight models are changing the cost math for anyone budgeting AI, and the bigger question every engineering leader is weighing — where "AI takes our jobs" actually leads if you follow it out.

...more
View all episodesView all episodes
Download on the App Store

The Forward Slash PodcastBy Callibrity