Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Bridging the AI–Data Gap: Collect, Curate, Serve
    Summary
    In this episode of the Data Engineering Podcast Omri Lifshitz (CTO) and Ido Bronstein (CEO) of Upriver talk about the growing gap between AI's demand for high-quality data and organizations' current data practices. They discuss why AI accelerates both the supply and demand sides of data, highlighting that the bottleneck lies in the "middle layer" of curation, semantics, and serving. Omri and Ido outline a three-part framework for making data usable by LLMs and agents: collect, curate, serve, and share challenges of scaling from POCs to production, including compounding error rates and reliability concerns. They also explore organizational shifts, patterns for managing context windows, pragmatic views on schema choices, and Upriver's approach to building autonomous data workflows using determinism and LLMs at the right boundaries. The conversation concludes with a look ahead to AI-first data platforms where engineers supervise business semantics while automation stitches technical details end-to-end.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Composable data infrastructure is great, until you spend all of your time gluing it together. Bruin is an open source framework, driven from the command line, that makes integration a breeze. Write Python and SQL to handle the business logic, and let Bruin handle the heavy lifting of data movement, lineage tracking, data quality monitoring, and governance enforcement. Bruin allows you to build end-to-end data workflows using AI, has connectors for hundreds of platforms, and helps data teams deliver faster. Teams that use Bruin need less engineering effort to process data and benefit from a fully integrated data platform. Go to dataengineeringpodcast.com/bruin today to get started. And for dbt Cloud customers, they'll give you $1,000 credit to migrate to Bruin Cloud.
    • Your host is Tobias Macey and today I'm interviewing Omri Lifshitz and Ido Bronstein about the challenges of keeping up with the demand for data when supporting AI systems
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • We're here to talk about "The Growing Gap Between Data & AI". From your perspective, what is this gap, and why do you think it's widening so rapidly right now?
    • How does this gap relate to the founding story of Upriver? What problems were you and your co-founders experiencing that led you to build this?
    • The core premise of new AI tools, from RAG pipelines to LLM agents, is that they are only as good as the data they're given. How does this "garbage in, garbage out" problem change when the "in" is not a static file but a complex, high-velocity, and constantly changing data pipeline?
    • Upriver is described as an "intelligent agent system" and an "autonomous data engineer." This is a fascinating "AI to solve for AI" approach. Can you describe this agent-based architecture and how it specifically works to bridge that data-AI gap?
    • Your website mentions a "Data Context Layer" that turns "tribal knowledge" into a "machine-usable mode." This sounds critical for AI. How do you capture that context, and how does it make data "AI-ready" in a way that a traditional data catalog or quality tool doesn't?
    • What are the most innovative or unexpected ways you've seen companies trying to make their data "AI-ready"? And where are the biggest points of failure you observe?
    • What has been the most challenging or unexpected lesson you've learned while building an AI system (Upriver) that is designed to fix the data foundation for other AI systems?
    • When is an autonomous, agent-based approach not the right solution for a team's data quality problems? What organizational or technical maturity is required to even start closing this data-AI gap?
    • What do you have planned for the future of Upriver? And looking more broadly, how do you see this gap between data and AI evolving over the next few years?
    Contact Info
    • Ido - LinkedIn
    • Omri - LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Upriver
    • RAG == Retrieval Augmented Generation
      • AI Engineering Podcast Episode
    • AI Agent
    • Context Window
    • Model Finetuning)
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    51 min
  • Beyond the Perimeter: Practical Patterns for Fine‑Grained Data Access
    Summary
    In this episode of the Data Engineering Podcast Matt Topper, president of UberEther, talks about the complex challenge of identity, credentials, and access control in modern data platforms. With the shift to composable ecosystems, integration burdens have exploded, fracturing governance and auditability across warehouses, lakes, files, vector stores, and streaming systems. Matt shares practical solutions, including propagating user identity via JWTs, externalizing policy with engines like OPA/Rego and Cedar, and using database proxies for native row/column security. He also explores catalog-driven governance, lineage-based label propagation, and OpenTDF for binding policies to data objects. The conversation covers machine-to-machine access, short-lived credentials, workload identity, and constraining access by interface choke points, as well as lessons from Zanzibar-style policy models and the human side of enforcement. Matt emphasizes the need for trust composition - unifying provenance, policy, and identity context - to answer questions about data access, usage, and intent across the entire data path.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Composable data infrastructure is great, until you spend all of your time gluing it together. Bruin is an open source framework, driven from the command line, that makes integration a breeze. Write Python and SQL to handle the business logic, and let Bruin handle the heavy lifting of data movement, lineage tracking, data quality monitoring, and governance enforcement. Bruin allows you to build end-to-end data workflows using AI, has connectors for hundreds of platforms, and helps data teams deliver faster. Teams that use Bruin need less engineering effort to process data and benefit from a fully integrated data platform. Go to dataengineeringpodcast.com/bruin today to get started. And for dbt Cloud customers, they'll give you $1,000 credit to migrate to Bruin Cloud.
    • Your host is Tobias Macey and today I'm interviewing Matt Topper about the challenges of managing identity and access controls in the context of data systems
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • The data ecosystem is a uniquely challenging space for creating and enforcing technical controls for identity and access control. What are the key considerations for designing a strategy for addressing those challenges?
    • For data acess the off-the-shelf options are typically on either extreme of too coarse or too granular in their capabilities. What do you see as the major factors that contribute to that situation?
    • Data governance policies are often used as the primary means of identifying what data can be accesssed by whom, but translating that into enforceable constraints is often left as a secondary exercise. How can we as an industry make that a more manageable and sustainable practice?
    • How can the audit trails that are generated by data systems be used to inform the technical controls for identity and access?
    • How can the foundational technologies of our data platforms be improved to make identity and authz a more composable primitive?
    • How does the introduction of streaming/real-time data ingest and delivery complicate the challenges of security controls?
    • What are the most interesting, innovative, or unexpected ways that you have seen data teams address ICAM?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on ICAM?
    • What are the aspects of ICAM in data systems that you are paying close attention to?
      • What are your predictions for the industry adoption or enforcement of those controls?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • UberEther
    • JWT == JSON Web Token
    • OPA == Open Policy Agent
    • Rego
    • PingIdentity
    • Okta
    • Microsoft Entra
    • SAML == Security Assertion Markup Language
    • OAuth
    • OIDC == OpenID Connect
    • IDP == Identity Provider
    • Kubernetes
    • Istio
    • Amazon CEDAR policy language
    • AWS IAM
    • PII == Personally Identifiable Information
    • CISO == Chief Information Security Officer
    • OpenTDF
    • OpenFGA
    • Google Zanzibar
    • Risk Management Framework
    • Model Context Protocol
    • Google Data Project
    • TPM == Trusted Platform Module
    • PKI == Public Key Infrastructure
    • Passskeys
    • DuckLake
      • Podcast Episode
    • Accumulo
    • JDBC
    • OpenBao
    • Hashicorp Vault
    • LDAP
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 5 min
  • The True Costs of Legacy Systems: Technical Debt, Risk, and Exit Strategies
    Summary
    In this episode Kate Shaw, Senior Product Manager for Data and SLIM at SnapLogic, talks about the hidden and compounding costs of maintaining legacy systems—and practical strategies for modernization. She unpacks how “legacy” is less about age and more about when a system becomes a risk: blocking innovation, consuming excess IT time, and creating opportunity costs. Kate explores technical debt, vendor lock-in, lost context from employee turnover, and the slippery notion of “if it ain’t broke,” especially when data correctness and lineage are unclear. Shee digs into governance, observability, and data quality as foundations for trustworthy analytics and AI, and why exit strategies for system retirement should be planned from day one. The discussion covers composable architectures to avoid monoliths and big-bang migrations, how to bridge valuable systems into AI initiatives without lock-in, and why clear success criteria matter for AI projects. Kate shares lessons from the field on discovery, documentation gaps, parallel run strategies, and using integration as the connective tissue to unlock data for modern, cloud-native and AI-enabled use cases. She closes with guidance on planning migrations, defining measurable outcomes, ensuring lineage and compliance, and building for swap-ability so teams can evolve systems incrementally instead of living with a “bowl of spaghetti.”

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Kate Shaw about the true costs of maintaining legacy systems
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • What are your crtieria for when a given system or service transitions to being "legacy"?
    • In order for any service to survive long enough to become "legacy" it must be serving its purpose and providing value. What are the common factors that prompt teams to deprecate or migrate systems?
    • What are the sources of monetary cost related to maintaining legacy systems while they remain operational?
    • Beyond monetary cost, economics also have a concept of "opportunity cost". What are some of the ways that manifests in data teams who are maintaining or migrating from legacy systems?
      • How does that loss of productivity impact the broader organization?
    • How does the process of migration contribute to issues around data accuracy, reliability, etc. as well as contributing to potential compromises of security and compliance?
    • Once a system has been replaced, it needs to be retired. What are some of the costs associated with removing a system from service?
    • What are the most interesting, innovative, or unexpected ways that you have seen teams address the costs of legacy systems and their retirement?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on legacy systems migration?
    • When is deprecation/migration the wrong choice?
    • How have evolutionary architecture patterns helped to mitigate the costs of system retirement?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • SnapLogic
    • SLIM == SnapLogic Intelligent Modernizer
    • Opportunity Cost
    • Sunk Cost Fallacy
    • Data Governance
    • Evolutionary Architecture
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 5 min
  • Context Engineering as a Discipline: Building Governed AI Analytics
    Summary
    In this episode of the Data Engineering Podcast, host Tobias Macey welcomes back Nick Schrock, CTO and founder of Dagster Labs, to discuss Compass - a Slack-native, agentic analytics system designed to keep data teams connected with business stakeholders. Nick shares his journey from initial skepticism to embracing agentic AI as model and application advancements made it practical for governed workflows, and explores how Compass redefines the relationship between data teams and stakeholders by shifting analysts into steward roles, capturing and governing context, and integrating with Slack where collaboration already happens. The conversation covers organizational observability through Compass's conversational system of record, cost control strategies, and the implications of agentic collaboration on Conway's Law, as well as what's next for Compass and Nick's optimistic views on AI-accelerated software engineering.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • Your host is Tobias Macey and today I'm interviewing Nick Schrock about building an AI analyst that keeps data teams in the loop
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Compass is and the story behind it?
    • context repository structure
      • how to keep it relevant/avoid sprawl/duplication
    • providing guardrails
    • how does a tool like Compass help provide feedback/insights back to the data teams?
    • preparing the data warehouse for effective introspection by the AI
    • LLM selection
    • cost management
      • caching/materializing ad-hoc queries
    • Why Slack and enterprise chat are important to b2b software
    • How AI is changing stakeholder relationships
    • How not to overpromise AI capabilities 
    • How does Compass relate to BI?
    • How does Compass relate to Dagster and Data Infrastructure?
    • What are the most interesting, innovative, or unexpected ways that you have seen Compass used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Compass?
    • When is Compass the wrong choice?
    • What do you have planned for the future of Compass?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Dagster
    • Dagster Labs
    • Dagster Plus
    • Dagster Compass
    • Chris Bergh DataOps Episode
    • Rise of Medium Code blog post
    • Context Engineering
    • Data Steward
    • Information Architecture
    • Conway's Law
    • Temporal durable execution framework
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    52 min
  • The Data Model That Captures Your Business: Metric Trees Explained
    Summary
    In this episode of the Data Engineering Podcast Vijay Subramanian, founder and CEO of Trace, talks about metric trees - a new approach to data modeling that directly captures a company's business model. Vijay shares insights from his decade-long experience building data practices at Rent the Runway and explains how the modern data stack has led to a proliferation of dashboards without a coherent way for business consumers to reason about cause, effect, and action. He explores how metric trees differ from and interoperate with other data modeling approaches, serve as a backend for analytical workflows, and provide concrete examples like modeling Uber's revenue drivers and customer journeys. Vijay also discusses the potential of AI agents operating on metric trees to execute workflows, organizational patterns for defining inputs and outputs with business teams, and a vision for analytics that becomes invisible infrastructure embedded in everyday decisions.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Vijay Subramanian about metric trees and how they empower more effective and adaptive analytics
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what metric trees are and their purpose?
    • How do metric trees relate to metric/semantic layers?
    • What are the shortcomings of existing data modeling frameworks that prevent effective use of those assets?
      • How do metric trees build on top of existing investments in dimensional data models?
    • What are some strategies for engaging with the business to identify metrics and their relationships?
    • What are your recommendations for storage, representation, and retrieval of metric trees?
    • How do metric trees fit into the overall lifecycle of organizational data workflows?
    • When creating any new data asset it introduces overhead of maintenance, monitoring, and evolution. How do metric trees fit into existing testing and validation frameworks that teams rely on for dimensional modeling?
      • What are some of the key differences in useful evaluation/testing that teams need to develop for metric trees?
    • How do metric trees assist in context engineering for AI-powered self-serve access to organizational data?
    • What are the most interesting, innovative, or unexpected ways that you have seen metric trees used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on metric trees and operationalizing them at Trace?
    • When is a metric tree the wrong abstraction?
    • What do you have planned for the future of Trace and applications of metric trees?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Metric Tree
    • Trace
    • Modern Data Stack
    • Hadoop
    • Vertica
    • Luigi
    • dbt
    • Ralph Kimball
    • Bill Inmon
    • Metric Layer
    • Dimensional Data Warehouse
    • Master Data Management
    • Data Governance
    • Financial P&L (Profit and Loss)
    • EBITDA ==Earnings before interest, taxes, depreciation and amortization
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 2 min
  • From GPUs-as-a-Service to Workloads-as-a-Service: Flex AI’s Path to High-Utilization AI Infra
    Summary
    In this crossover episode of the AI Engineering Podcast, host Tobias Macey interviews Brijesh Tripathi, CEO of Flex AI, about revolutionizing AI engineering by removing DevOps burdens through "workload as a service". Brijesh shares his expertise from leading AI/HPC architecture at Intel and deploying supercomputers like Aurora, highlighting how access friction and idle infrastructure slow progress. Join them as they discuss Flex AI's innovative approach to simplifying heterogeneous compute, standardizing on consistent Kubernetes layers, and abstracting inference across various accelerators, allowing teams to iterate faster without wrestling with drivers, libraries, or cloud-by-cloud differences. Brijesh also shares insights into Flex AI's strategies for lifting utilization, protecting real-time workloads, and spanning the full lifecycle from fine-tuning to autoscaled inference, all while keeping complexity at bay.

    Pre-amble
    I hope you enjoy this cross-over episode of the AI Engineering Podcast, another show that I run to act as your guide to the fast-moving world of building scalable and maintainable AI systems. As generative AI models have grown more powerful and are being applied to a broader range of use cases, the lines between data and AI engineering are becoming increasingly blurry. The responsibilities of data teams are being extended into the realm of context engineering, as well as designing and supporting new infrastructure elements that serve the needs of agentic applications. This episode is an example of the types of work that are not easily categorized into one or the other camp.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • Your host is Tobias Macey and today I'm interviewing Brijesh Tripathi about FlexAI, a platform offering a service-oriented abstraction for AI workloads
    Interview
    • Introduction
    • How did you get involved in machine learning?
    • Can you describe what FlexAI is and the story behind it?
    • What are some examples of the ways that infrastructure challenges contribute to friction in developing and operating AI applications?
      • How do those challenges contribute to issues when scaling new applications/businesses that are founded on AI?
    • There are numerous managed services and deployable operational elements for operationalizing AI systems. What are some of the main pitfalls that teams need to be aware of when determining how much of that infrastructure to own themselves?
    • Orchestration is a key element of managing the data and model lifecycles of these applications. How does your approach of "workload as a service" help to mitigate some of the complexities in the overall maintenance of that workload?
    • Can you describe the design and architecture of the FlexAI platform?
      • How has the implementation evolved from when you first started working on it?
    • For someone who is going to build on top of FlexAI, what are the primary interfaces and concepts that they need to be aware of?
    • Can you describe the workflow of going from problem to deployment for an AI workload using FlexAI?
    • One of the perennial challenges of making a well-integrated platform is that there are inevitably pre-existing workloads that don't map cleanly onto the assumptions of the vendor. What are the affordances and escape hatches that you have built in to allow partial/incremental adoption of your service?
    • What are the elements of AI workloads and applications that you are explicitly not trying to solve for?
    • What are the most interesting, innovative, or unexpected ways that you have seen FlexAI used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on FlexAI?
    • When is FlexAI the wrong choice?
    • What do you have planned for the future of FlexAI?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what are the biggest gaps in tooling, technology, or training for AI systems today?
    Links
    • Flex AI
    • Aurora Super Computer
    • CoreWeave
    • Kubernetes
    • CUDA
    • ROCm
    • Tensor Processing Unit (TPU)
    • PyTorch
    • Triton
    • Trainium
    • ASIC == Application Specific Integrated Circuit
    • SOC == System On a Chip
    • Loveable
    • FlexAI Blueprints
    • Tenstorrent
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    57 min
  • From RAG to Relational: How Agentic Patterns Are Reshaping Data Architecture
    Summary
    In this episode of the AI Engineering Podcast Mark Brooker, VP and Distinguished Engineer at AWS, talks about how agentic workflows are transforming database usage and infrastructure design. He discusses the evolving role of data in AI systems, from traditional models to more modern approaches like vectors, RAG, and relational databases. Mark explains why agents require serverless, elastic, and operationally simple databases, and how AWS solutions like Aurora and DSQL address these needs with features such as rapid provisioning, automated patching, geodistribution, and spiky usage. The conversation covers topics including tool calling, improved model capabilities, state in agents versus stateless LLM calls, and the role of Lambda and AgentCore for long-running, session-isolated agents. Mark also touches on the shift from local MCP tools to secure, remote endpoints, the rise of object storage as a durable backplane, and the need for better identity and authorization models. The episode highlights real-world patterns like agent-driven SQL fuzzing and plan analysis, while identifying gaps in simplifying data access, hardening ops for autonomous systems, and evolving serverless database ergonomics to keep pace with agentic development.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Marc Brooker about the impact of agentic workflows on database usage patterns and how they change the architectural requirements for databases
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what the role of the database is in agentic workflows?
      • There are numerous types of databases, with relational being the most prevalent. How does the type and purpose of an agent inform the type of database that should be used?
    • Anecdotally I have heard about how agentic workloads have become the predominant "customers" of services like Neon and Fly.io. How would you characterize the different patterns of scale for agentic AI applications? (e.g. proliferation of agents, monolithic agents, multi-agent, etc.)
    • What are some of the most significant impacts on workload and access patterns for data storage and retrieval that agents introduce?
      • What are the categorical differences in that behavior as compared to programmatic/automated systems?
    • You have spent a substantial amount of time on Lambda at AWS. Given that LLMs are effectively stateless, how does the added ephemerality of serverless functions impact design and performance considerations around having to "re-hydrate" context when interacting with agents?
    • What are the most interesting, innovative, or unexpected ways that you have seen serverless and database systems used for agentic workloads?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on technologies that are supporting agentic applications?
    Contact Info
    • Blog
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • AWS Aurora DSQL
    • AWS Lambda
    • Three Tier Architecture
    • Vector Database
    • Graph Database
    • Relational Database
    • Vector Embedding
    • RAG == Retrieval Augmented Generation
      • AI Engineering Podcast Episode
    • GraphRAG
      • AI Engineering Podcast Episode
    • LLM Tool Calling
    • MCP == Model Context Protocol
    • A2A == Agent 2 Agent Protocol
    • AWS Bedrock AgentCore
    • Strands
    • LangChain
    • Kiro
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    53 min
  • Duck Lake: Simplifying the Lakehouse Ecosystem
    Summary
    In this episode of the Data Engineering Podcast Hannes Mühleisen and Mark Raasveldt, the creators of DuckDB, share their work on Duck Lake, a new entrant in the open lakehouse ecosystem. They discuss how Duck Lake, is focused on simplicity, flexibility, and offers a unified catalog and table format compared to other lakehouse formats like Iceberg and Delta. Hannes and Mark share insights into how Duck Lake revolutionizes data architecture by enabling local-first data processing, simplifying deployment of lakehouse solutions, and offering benefits such as encryption features, data inlining, and integration with existing ecosystems.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data teams everywhere face the same problem: they're forcing ML models, streaming data, and real-time processing through orchestration tools built for simple ETL. The result? Inflexible infrastructure that can't adapt to different workloads. That's why Cash App and Cisco rely on Prefect. Cash App's fraud detection team got what they needed - flexible compute options, isolated environments for custom packages, and seamless data exchange between workflows. Each model runs on the right infrastructure, whether that's high-memory machines or distributed compute. Orchestration is the foundation that determines whether your data team ships or struggles. ETL, ML model training, AI Engineering, Streaming - Prefect runs it all from ingestion to activation in one platform. Whoop and 1Password also trust Prefect for their data operations. If these industry leaders use Prefect for critical workflows, see what it can do for you at dataengineeringpodcast.com/prefect.
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • Your host is Tobias Macey and today I'm interviewing Hannes Mühleisen and Mark Raasveldt about DuckLake, the latest entrant into the open lakehouse ecosystem
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what DuckLake is and the story behind it?
      • What are the particular problems that DuckLake is solving for?
    • How does this compare to the capabilities of MotherDuck?
    • Iceberg and Delta already have a well established ecosystem, but so does DuckDB. Who are the primary personas that you are trying to focus on in these early days of DuckLake?
    • One of the major factors driving the adoption of formats like Iceberg is cost efficiency for large volumes of data. That brings with it challenges of large batch processing of data. How does DuckLake account for these axes of scale?
    • There is also a substantial investment in the ecosystem of technologies that support Iceberg. The most notable ecosystem challenge for DuckDB and DuckLake is in the query layer. How are you thinking about the evolution and growth of that capability beyond DuckDB (e.g. support in Trino/Spark/Flink)?
    • What are your opinions on the viability of a future where DuckLake and Iceberg become a unified standard and implementation? (why can't Iceberg REST catalog implementations just use DuckLake under the hood?)
    • Digging into the specifics of the specification and implementation, what are some of the capabilities that it offers above and beyond Iceberg?
      • Is it now possible to enforce PK/FK constraints, indexing on underlying data?
    • Given that DuckDB has a vector type, how do you think about the support for vector storage/indexing?
    • How do the capabilities of DuckLake and the integration with DuckDB change the ways that data teams design their data architecture and access patterns?
    • What are your thoughts on the impact of "data gravity" in today's data ecosystem, with engines like DuckDB, KuzuDB, LanceDB, etc. available for embedded and edge use cases?
    • What are the most interesting, innovative, or unexpected ways that you have seen DuckLake used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on DuckLake?
    • When is DuckLake the wrong choice?
    • What do you have planned for the future of DuckLake?
    Contact Info
    • Hannes
      • Website
    • Mark
      • Website
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • DuckDB
      • Podcast Episode
    • DuckLake
    • DuckDB Labs
    • MySQL
    • CWI
    • MonetDB
    • Iceberg
    • Iceberg REST Catalog
    • Delta
    • Hudi
    • Lance
    • DuckDB Iceberg Connector
    • ACID == Atomicity, Consistency, Isolation, Durability
    • MotherDuck
    • MotherDuck Managed DuckLake
    • Trino
    • Spark
    • Presto
    • Spark DuckLake Demo
    • Delta Kernel
    • Arrow
    • dlt
    • S3 Tables
    • Attribute Based Access Control (ABAC)
    • Parquet
    • Arrow Flight
    • Hadoop
    • HDFS
    • DuckLake Roadmap
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 11 min
  • Aligning Business and Data: The Essential Role of Data Modeling
    Summary
    In this episode of the Data Engineering Podcast Serge Gershkovich, head of product at SQL DBM, talks about the socio-technical aspects of data modeling. Serge shares his background in data modeling and highlights its importance as a collaborative process between business stakeholders and data teams. He debunks common misconceptions that data modeling is optional or secondary, emphasizing its crucial role in ensuring alignment between business requirements and data structures. The conversation covers challenges in complex environments, the impact of technical decisions on data strategy, and the evolving role of AI in data management. Serge stresses the need for business stakeholders' involvement in data initiatives and a systematic approach to data modeling, warning against relying solely on technical expertise without considering business alignment.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Enterprises today face an enormous challenge: they’re investing billions into Snowflake and Databricks, but without strong foundations, those investments risk becoming fragmented, expensive, and hard to govern. And that’s especially evident in large, complex enterprise data environments. That’s why companies like DirecTV and Pfizer rely on SqlDBM. Data modeling may be one of the most traditional practices in IT, but it remains the backbone of enterprise data strategy. In today’s cloud era, that backbone needs a modern approach built natively for the cloud, with direct connections to the very platforms driving your business forward. Without strong modeling, data management becomes chaotic, analytics lose trust, and AI initiatives fail to scale. SqlDBM ensures enterprises don’t just move to the cloud—they maximize their ROI by creating governed, scalable, and business-aligned data environments. If global enterprises are using SqlDBM to tackle the biggest challenges in data management, analytics, and AI, isn’t it worth exploring what it can do for yours? Visit dataengineeringpodcast.com/sqldbm to learn more.
    • Your host is Tobias Macey and today I'm interviewing Serge Gershkovich about how and why data modeling is a sociotechnical endeavor
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing the activities that you think of when someone says the term "data modeling"?
      • What are the main groupings of incomplete or inaccurate definitions that you typically encounter in conversation on the topic?
      • How do those conceptions of the problem lead to challenges and bottlenecks in execution?
    • Data modeling is often associated with data warehouse design, but it also extends to source systems and unstructured/semi-structured assets. How does the inclusion of other data localities help in the overall success of a data/domain modeling effort?
    • Another aspect of data modeling that often consumes a substantial amount of debate is which pattern to adhere to (star/snowflake, data vault, one big table, anchor modeling, etc.). What are some of the ways that you have found effective to remove that as a stumbling block when first developing an organizational domain representation?
    • While the overall purpose of data modeling is to provide a digital representation of the business processes, there are inevitable technical decisions to be made. What are the most significant ways that the underlying technical systems can help or hinder the goals of building a digital twin of the business?
    • What impact (positive and negative) are you seeing from the introduction of LLMs into the workflow of data modeling?
      • How does tool use (e.g. MCP connection to warehouse/lakehouse) help when developing the transformation logic for achieving a given domain representation? 
    • What are the most interesting, innovative, or unexpected ways that you have seen organizations address the data modeling lifecycle?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working with organizations implementing a data modeling effort?
    • What are the overall trends in the ecosystem that you are monitoring related to data modeling practices?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Links
    • sqlDBM
    • SAP
    • Joe Reis
    • ERD == Entity Relation Diagram
    • Master Data Management
    • dbt
    • Data Contracts
    • Data Modeling With Snowflake book by Serge (affiliate link)
    • Type 2 Dimension
    • Data Vault
    • Star Schema
    • Anchor Modeling
    • Ralph Kimball
    • Bill Inmon
    • Sixth Normal Form
    • MCP == Model Context Protocol
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 7 min
  • From Academia to Industry: Bridging Data Engineering Challenges
    Summary
    In this episode of the Data Engineering Podcast Professor Paul Groth, from the University of Amsterdam, talks about his research on knowledge graphs and data engineering. Paul shares his background in AI and data management, discussing the evolution of data provenance and lineage, as well as the challenges of data integration. He explores the impact of large language models (LLMs) on data engineering, highlighting their potential to simplify knowledge graph construction and enhance data integration. The conversation covers the evolving landscape of data architectures, managing semantics and access control, and the interplay between industry and academia in advancing data engineering practices, with Paul also sharing insights into his work with the intelligent data engineering lab and the importance of human-AI collaboration in data engineering pipelines.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Paul Groth about his research on knowledge graphs and data engineering
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing the focus and scope of your academic efforts?
    • Given your focus on data management for machine learning as part of the INDELab, what are some of the developing trends that practitioners should be aware of?
      • ML architectures / systems changing (matteo interlandi) GPUs for data mangement
    • You have spent a large portion of your career working with knowledge graphs, which have largely been a niche area until recently. What are some of the notable changes in the knowledge graph ecosystem that have resulted from the introduction of LLMs?
    • What are some of the other ways that you are seeing LLMs change the methods of data engineering?
      • There are numerous vague and anecdotal references to the power of LLMs to unlock value from unstructured data. What are some of the realitites that you are seeing in your research?
    • A majority of the conversations in this podcast are focused on data engineering in the context of a business organization. What are some of the ways that management of research data is disjoint from the methods and constraints that are present in business contexts?
    • What are the most interesting, innovative, or unexpected ways that you have seen LLM used in data management?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data engineering research?
    • What do you have planned for the future of your research in the context of data engineering, knowledge graphs, and AI?
    Contact Info
    • Website
    • email
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • INDELab
    • Data Provenance
    • Elsevier
    • SIGMOD 2025
    • Digital Twin
    • Knowledge Graph
    • WikiData
    • KuzuDB
      • Podcast Episode
    • data.world
      • Podcast Episode
    • GraphRAG
    • SPARQL
    • Semantic Web
    • GQL == Graph Query Language
    • Cypher
    • Amazon Neptune
    • RDF == Resource Description Framework
    • SwellDB
    • FlockMTL
    • DuckDB
      • Podcast Episode
    • Matteo Interlandi
    • Paolo Papotti
    • Neuromorphic Computing
    • Point Clouds
    • Longform.ai
    • BASIL DB
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    51 min

About Data Engineering Podcast

From the publisher's feed

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

More shows like Data Engineering Podcast

This Week in Startups by Jason Calacanis

This Week in Startups

1,290 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,087 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Risky Business by Risky Business Media

Risky Business

375 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

582 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

Syntax - Tasty Web Development Treats

985 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

The Data Engineering Show by The Firebolt Data Bros

The Data Engineering Show

8 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

102 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

684 Listeners