Coinbase drove its token usage up while cutting inference costs by more than half. Kyle Cesmat, who runs Coinbase's Agent Experience team, explains the architecture behind it, from an internal LLM gateway that routes their inference to cache hit rates held above 85 percent.
He walks Ameya through his goal of moving 60 percent of engineering work to cloud agents, treating them as auditable service accounts inside a regulated company, and an approval engine that merges some changes without a human. His advice for smaller teams is to narrow their tooling and own inference observability.
Topics discussed:
- Routing controlled inference through an internal LLM gateway
- Holding cache hit rates above 85 percent
- Why changing models mid-session busts your prompt cache
- Moving 60 percent of engineering work to cloud agents
- Treating agents as auditable service accounts with strong identity
- Running an approval engine that merges without human review
- Setting token budgets for both developers and agents
- Narrowing your tooling aperture to control AI costs
Listen to more episodes:
Apple
Spotify
YouTube