
Sign up to save your podcasts
Or


Today we are talking with Vignesh Baskaran, the CTO and co-founder of Hexo Labs, about teaching AI agents to improve themselves. Vignesh has been training neural networks since 2012, back when he was still called a data scientist.
Then he became an ML engineer and now an AI engineer, though he says the underlying work has never really changed. It's to figure out how to make a system behave the way you intend it to. He built the litigation search engine that Google itself became a customer of, and now he's chasing something new, agents that rewrite and retrain other agents without a human in the loop.
We dig into Sia, the meta-agent at the center of Hexo's research, and why improving an agent means touching both its harness and its actual model weights, not just one or the other. We talk about proxy evals for when you don't have much to ground truth. The Darwin-Gödel machine and why formal verification is too strict a bar for anything commercial.
How Hexo's work echoes DeepMind's Alpha lineage from AlphaGo to AlphaEvolve, and the spectrum from clearly verifiable to totally subjective tasks? Why VAE evals are quietly wrecking agent quality across the industry, and a great story about an agent that discovered a customer's own eval file was silently corrupted, something buried in hundreds of thousands of traces that no human would have caught.
Today we're talking with Kolton Andrus, the Founder and CEO of Gremlin, about what happens to reliability when AI is writing most of the code. Kolton helped build the Chaos Engineering practice of both Amazon and Netflix before starting Gremlin.
In our conversation we talk about scar tissue, the intuition engineers develop from being woken up at 3:00 AM to fix production outages and how AI doesn't have any of it. It generates code in an afternoon that maybe took a team previously weeks to build, but none of those painful lessons come along for the ride.
We dig into why 10x more code might mean 10x more failures. The concept of reliability guardrails, think ethical guardrails, but for keeping your systems up. Why you still have to test in production no matter how good your staging environment is? How Gremlin is rethinking their product for the world where agents, not engineers, are essentially the primary users.And why we're entering a painful, narrow part of the hourglass before AI gets good enough to handle all of this on its own.
Today we are talking with Andre Elizondo, the Director of Innovation at Mezmo about their open source agentic harness for SREs called AURA. Mezmo got their start handling observability data at scale. Logs, traces, metrics, the usual stuff.
Today on the show, we have a special guest — Ashmeet Sidana, the founder of Engineering Capital.
Today we have Dr. Ewelina Kurtys on the show. Ewelina has a background in Neuroscience and is currently working at FinalSpark.
FinalSpark is using live Neurons for computations instead of traditional electric CPUs. The advantage is that live Neurons are significantly more energy efficient than traditional computing, and given all the energy concerns right now with regards to running AI workloads and data centers, this seems quite relevant, even though bioprocessors are still very much in the research phase.
Today we're talking with one of our favorite engineers, Rafal Wilinski. Rafal has been on the cutting edge of AI development in the last few years as he has led AI teams at Zapier and Vendr.
If you’ve ever felt like engineering teams are stuck in execution mode—heads down, building what they’re told—then today’s episode is for you. We're talking about what it really takes to build high ownership engineering cultures where devs aren't simply just shipping code, but they're helping shape the product.
And our guest this week is Matt Watson. He's a long time founder, engineer, and now the CEO of Full Scale, a company that helps startups and scale ups, grow their engineering teams with top talent from the Philippines. Matt's also the author of a book called Product Driven that shows how engineers can build with more clarity, purpose, customer focus and we get into some of the details in that book during this podcast.
So in this episode, we get into everything from the downsides of specialization to the importance of empathy, to why code shipped isn't the same as value delivered. We hope you enjoy it.
Today's episode is with Aayush Shah. Aayush is one of the co-founders of Blacksmith, which is a CI compute platform. Basically, Blacksmith will run your GitHub Actions jobs faster and with more visibility with the standard GitHub Actions CI runners.
The founding team has a fun background doing systems work at Cockroach and Faire, and they're taking on a big problem in running this massive CI fleet.
The explosion in AI agents has really changed the CI world. CI is more useful than ever, as you want to be sure the changes from your agents aren't breaking your existing functionality. At the same time, there's a huge increase in demand and spikiness of CI workloads as developers can fire off multiple agents to work in parallel, each needing to run the CI suite before merging. Aayush talked about how they're handling this load and facilitating visibility into test failures.
We also covered cloud economics. Aayush said the traditional cloud-based storage options don't work for them -- EBS and locally attached SSDs are too expensive for their workloads where they don't need the standard durability guarantees. He walks us through building their own fleet outside the hyperscalers and the plans going forward, along with some of the economics of multi-tenancy that Blacksmith has previously written about.
Today, we're talking Valkey, Redis, and all things caching. Our guest is Madelyn Olson, who is a principal engineer at AWS working on Elasticache and is one of the most well-known people in the caching community. She was a core maintainer of Redis prior to the fork and was one of the creators of Valkey, an open-source fork of Redis.
Today, Sam Lambert from Planetscale is back for a third time. Planetscale just announced Planetscale Postgres, so we had to get Sam back to tell us how and why they decided to add support for Postgres. It's always great to have Sam on -- he brings great stories about real customers and honest insight about the state of the database industry.
From the publisher's feed

286 Listeners

66 Listeners

4 Listeners