Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Amazon S3: The Backbone of Modern Data Systems
    Summary
    In this episode of the Data Engineering Podcast Mai-Lan Tomsen Bukovec, Vice President of Technology at AWS, talks about the evolution of Amazon S3 and its profound impact on data architecture. From her work on compute systems to leading the development and operations of S3, Mylan shares insights on how S3 has become a foundational element in modern data systems, enabling scalable and cost-effective data lakes since its launch alongside Hadoop in 2006. She discusses the architectural patterns enabled by S3, the importance of metadata in data management, and how S3's evolution has been driven by customer needs, leading to innovations like strong consistency and S3 tables.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • This is a pharmaceutical Ad for Soda Data Quality. Do you suffer from chronic dashboard distrust? Are broken pipelines and silent schema changes wreaking havoc on your analytics? You may be experiencing symptoms of Undiagnosed Data Quality Syndrome — also known as UDQS. Ask your data team about Soda. With Soda Metrics Observability, you can track the health of your KPIs and metrics across the business — automatically detecting anomalies before your CEO does. It’s 70% more accurate than industry benchmarks, and the fastest in the category, analyzing 1.1 billion rows in just 64 seconds. And with Collaborative Data Contracts, engineers and business can finally agree on what “done” looks like — so you can stop fighting over column names, and start trusting your data again.Whether you’re a data engineer, analytics lead, or just someone who cries when a dashboard flatlines, Soda may be right for you. Side effects of implementing Soda may include: Increased trust in your metrics, reduced late-night Slack emergencies, spontaneous high-fives across departments, fewer meetings and less back-and-forth with business stakeholders, and in rare cases, a newfound love of data. Sign up today to get a chance to win a $1000+ custom mechanical keyboard. Visit dataengineeringpodcast.com/soda to sign up and follow Soda’s launch week. It starts June 9th.
    • Your host is Tobias Macey and today I'm interviewing Mai-Lan Tomsen Bukovec about the evolutions of S3 and how it has transformed data architecture
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Most everyone listening knows what S3 is, but can you start by giving a quick summary of what roles it plays in the data ecosystem?
    • What are the major generational epochs in S3, with a particular focus on analytical/ML data systems?
      • The first major driver of analytical usage for S3 was the Hadoop ecosystem. What are the other elements of the data ecosystem that helped shape the product direction of S3?
    • Data storage and retrieval have been core primitives in computing since its inception. What are the characteristics of S3 and all of its copycats that led to such a difference in architectural patterns vs. other shared data technologies? (e.g. NFS, Gluster, Ceph, Samba, etc.)
    • How does the unified pool of storage that is exemplified by S3 help to blur the boundaries between application data, analytical data, and ML/AI data?
    • What are some of the default patterns for storage and retrieval across those three buckets that can lead to anti-patterns which add friction when trying to unify those use cases?
    • The age of AI is leading to a massive potential for unlocking unstructured data, for which S3 has been a massive dumping ground over the years. How is that changing the ways that your customers think about the value of the assets that they have been hoarding for so long?
      • What new architectural patterns is that generating?
    • What are the most interesting, innovative, or unexpected ways that you have seen S3 used for analytical/ML/Ai applications?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on S3?
    • When is S3 the wrong choice?
    • What do you have planned for the future of S3?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • AWS S3
    • Kinesis
    • Kafka
    • SQS
    • EMR
    • Drupal
    • Wordpress
    • Netflix Blog on S3 as a Source of Truth
    • Hadoop
    • MapReduce
    • Nasa JPL
    • FINRA == Financial Industry Regulatory Authority
    • S3 Object Versioning
    • S3 Cross Region
    • S3 Tables
    • Iceberg
    • Parquet
    • AWS KMS
    • Iceberg REST
    • DuckDB
    • NFS == Network File System
    • Samba
    • GlusterFS
    • Ceph
    • MinIO
    • S3 Metadata
    • Photoshop Generative Fill
    • Adobe Firefly
    • Turbotax AI Assistant
    • AWS Access Analyzer
    • Data Products
    • S3 Access Point
    • AWS Nova Models
    • LexisNexis Protege
    • S3 Intelligent Tiering
    • S3 Principal Engineering Tenets
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 2 min
  • Scaling Data Operations With Platform Engineering
    Summary
    In this episode of the Data Engineering Podcast Chakravarthy Kotaru talks about scaling data operations through standardized platform offerings. From his roots as an Oracle developer to leading the data platform at a major online travel company, Chakravarthy shares insights on managing diverse database technologies and providing databases as a service to streamline operations. He explains how his team has transitioned from DevOps to a platform engineering approach, centralizing expertise and automating repetitive tasks with AWS Service Catalog. Join them as they discuss the challenges of migrating legacy systems, integrating AI and ML for automation, and the importance of organizational buy-in in driving data platform success.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • This is a pharmaceutical Ad for Soda Data Quality. Do you suffer from chronic dashboard distrust? Are broken pipelines and silent schema changes wreaking havoc on your analytics? You may be experiencing symptoms of Undiagnosed Data Quality Syndrome — also known as UDQS. Ask your data team about Soda. With Soda Metrics Observability, you can track the health of your KPIs and metrics across the business — automatically detecting anomalies before your CEO does. It’s 70% more accurate than industry benchmarks, and the fastest in the category, analyzing 1.1 billion rows in just 64 seconds. And with Collaborative Data Contracts, engineers and business can finally agree on what “done” looks like — so you can stop fighting over column names, and start trusting your data again.Whether you’re a data engineer, analytics lead, or just someone who cries when a dashboard flatlines, Soda may be right for you. Side effects of implementing Soda may include: Increased trust in your metrics, reduced late-night Slack emergencies, spontaneous high-fives across departments, fewer meetings and less back-and-forth with business stakeholders, and in rare cases, a newfound love of data. Sign up today to get a chance to win a $1000+ custom mechanical keyboard. Visit dataengineeringpodcast.com/soda to sign up and follow Soda’s launch week. It starts June 9th.
    • Your host is Tobias Macey and today I'm interviewing Chakri Kotaru about scaling successful data operations through standardized platform offerings
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by outlining the different ways that you have seen teams you work with fail due to lack of structure and opinionated design?
    • Why NoSQL?
    • Pairing different styles of NoSQL for different problems
    • Useful patterns for each NoSQL style (document, column family, graph, etc.)
    • Challenges in platform automation and scaling edge cases
    • What challenges do you anticipate as a result of the new pressures as a result of AI applications?
    • What are the most interesting, innovative, or unexpected ways that you have seen platform engineering practices applied to data systems?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data platform engineering?
    • When is NoSQL the wrong choice?
    • What do you have planned for the future of platform principles for enabling data teams/data applications?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Riak
    • DynamoDB
    • SQL Server
    • Cassandra
    • ScyllaDB
    • CAP Theorem
    • Terraform
    • AWS Service Catalog
    • Blog Post
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    43 min
  • From Data Discovery to AI: The Evolution of Semantic Layers
    Summary
    In this episode of the Data Engineering Podcast, host Tobias Macy welcomes back Shinji Kim to discuss the evolving role of semantic layers in the era of AI. As they explore the challenges of managing vast data ecosystems and providing context to data users, they delve into the significance of semantic layers for AI applications. They dive into the nuances of semantic modeling, the impact of AI on data accessibility, and the importance of business logic in semantic models. Shinji shares her insights on how SelectStar is helping teams navigate these complexities, and together they cover the future of semantic modeling as a native construct in data systems. Join them for an in-depth conversation on the evolving landscape of data engineering and its intersection with AI.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Shinji Kim about the role of semantic layers in the era of AI
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Semantic modeling gained a lot of attention ~4-5 years ago in the context of the "modern data stack". What is your motivation for revisiting that topic today?
    • There are several overlapping concepts – "semantic layer," "metrics layer," "headless BI." How do you define these terms, and what are the key distinctions and overlaps?
      • Do you see these concepts converging, or do they serve distinct long-term purposes?
    • Data warehousing and business intelligence have been around for decades now. What new value does semantic modeling beyond practices like star schemas, OLAP cubes, etc.?
    • What benefits does a semantic model provide when integrating your data platform into AI use cases?
      • How is it different between using AI as an interface to your analytical use cases vs. powering customer facing AI applications with your data?
    • Putting in the effort to create and maintain a set of semantic models is non-zero. What role can LLMs play in helping to propose and construct those models?
      • For teams who have already invested in building this capability, what additional context and metadata is necessary to provide guidance to LLMs when working with their models?
    • What's the most effective way to create a semantic layer without turning it into a massive project? 
    • There are several technologies available for building and serving these models. What are the selection criteria that you recommend for teams who are starting down this path?
    • What are the most interesting, innovative, or unexpected ways that you have seen semantic models used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working with semantic modeling?
    • When is semantic modeling the wrong choice?
    • What do you predict for the future of semantic modeling?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • SelectStar
    • Sun Microsystems
    • Markov Chain Monte Carlo
    • Semantic Modeling
    • Semantic Layer
    • Metrics Layer
    • Headless BI
    • Cube
      • Podcast Episode
    • AtScale
    • Star Schema
    • Data Vault
    • OLAP Cube
    • RAG == Retrieval Augmented Generation
      • AI Engineering Podcast Episode
    • KNN == K-Nearest Neighbers
    • HNSW == Hierarchical Navigable Small World
    • dbt Metrics Layer
    • Soda Data
    • LookML
    • Hex
    • PowerBI
    • Tableau
    • Semantic View (Snowflake)
    • Databricks Genie
    • Snowflake Cortex Analyst
    • Malloy
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    50 min
  • Balancing Off-the-Shelf and Custom Solutions in Data Engineering
    Summary
    In this episode of the Data Engineering Podcast Tulika Bhatt, a senior software engineer at Netflix, talks about her experiences with large-scale data processing and the future of data engineering technologies. Tulika shares her journey into the data engineering field, discussing her work at BlackRock and Verizon before joining Netflix, and explains the challenges and innovations involved in managing Netflix's impression data for personalization and user experience. She highlights the importance of balancing off-the-shelf solutions with custom-built systems using technologies like Spark, Flink, and Iceberg, and delves into the complexities of ensuring data quality and observability in high-speed environments, including robust alerting strategies and semantic data auditing.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Tulika Bhatt about her experiences working on large scale data processing and her insights on the future trajectory of the supporting technologies
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by outlining the ways that operating at large scale change the ways that you need to think about the design of data systems?
    • When dealing with small-scale data systems it can be feasible to have manual processes. What are the elements of large scal data systems that demand autopmation?
      • How can those large-scale automation principles be down-scaled to the systems that the rest of the world are operating?
    • A perennial problem in data engineering is that of data quality. The past 4 years has seen a significant growth in the number of tools and practices available for automating the validation and verification of data. In your experience working with high volume data flows, what are the elements of data validation that are still unsolved?
    • Generative AI has taken the world by storm over the past couple years. How has that changed the ways that you approach your daily work?
    • What do you see as the future realities of working with data across various axes of large scale, real-time, etc.?
    • What are the most interesting, innovative, or unexpected ways that you have seen solutions to large-scale data management designed?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data management across axes of scale?
    • What are the ways that you are thinking about the future trajectory of your work??
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • BlackRock
    • Spark
    • Flink
    • Kafka
    • Cassandra
    • RocksDB
    • Netflix Maestro workflow orchestrator
    • Pagerduty
    • Iceberg
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    47 min
  • StarRocks: Bridging Lakehouse and OLAP for High-Performance Analytics
    Summary
    In this episode of the Data Engineering Podcast Sida Shen, product manager at CelerData, talks about StarRocks, a high-performance analytical database. Sida discusses the inception of StarRocks, which was forked from Apache Doris in 2020 and evolved into a high-performance Lakehouse query engine. He explains the architectural design of StarRocks, highlighting its capabilities in handling high concurrency and low latency queries, and its integration with open table formats like Apache Iceberg, Delta Lake, and Apache Hudi. Sida also discusses how StarRocks differentiates itself from other query engines by supporting on-the-fly joins and eliminating the need for denormalization pipelines, and shares insights into its use cases, such as customer-facing analytics and real-time data processing, as well as future directions for the platform.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Sida Shen about StarRocks, a high performance analytical database supporting shared nothing and shared data patterns
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what StarRocks is and the story behind it?
    • There are numerous analytical databases on the market. What are the attributes of StarRocks that differentiate it from other options?
    • Can you describe the architecture of StarRocks?
      • What are the "-ilities" that are foundational to the design of the system?
    • How have the design and focus of the project evolved since it was first created?
    • What are the tradeoffs involved in separating the communication layer from the data layers?
    • The tiered architecture enables the shared nothing and shared data behaviors, which allows for the implementation of lakehouse patterns. What are some of the patterns that are possible due to the single interface/dual pattern nature of StarRocks?
      • The shared data implementation has cacheing built in to accelerate interaction with datasets. What are some of the limitations/edge cases that operators and consumers should be aware of?
    • StarRocks supports management of lakehouse tables (Iceberg, Delta, Hudi, etc.), which overlaps with use cases for Trino/Presto/Dremio/etc. What are the cases where StarRocks acts as a replacement for those systems vs. a supplement to them?
    • The other major category of engines that StarRocks overlaps with is OLAP databases (e.g. Clickhouse, Firebolt, etc.). Why might someone use StarRocks in addition to or in place of those techologies?
    • We would be remiss if we ignored the dominating trend of AI and the systems that support it. What is the role of StarRocks in the context of an AI application?
    • What are the most interesting, innovative, or unexpected ways that you have seen StarRocks used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on StarRocks?
    • When is StarRocks the wrong choice?
    • What do you have planned for the future of StarRocks?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • StarRocks
    • CelerData
    • Apache Doris
    • SIMD == Single Instruction Multiple Data
    • Apache Iceberg
    • ClickHouse
      • Podcast Episode
    • Druid
    • Firebolt
      • Podcast Episode
    • Snowflake
    • BigQuery
    • Trino
    • Databricks
    • Dremio
    • Data Lakehouse
    • Delta Lake
    • Apache Hive
    • C++
    • Cost-Based Optimizer
    • Iceberg Summit Tencent Games Presentation
    • Apache Paimon
    • Lance
      • Podcast Episode
    • Delta Uniform
    • Apache Arrow
    • StarRocks Python UDF
    • Debezium
      • Podcast Episode
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr
  • Exploring NATS: A Multi-Paradigm Connectivity Layer for Distributed Applications
    Summary
    In this episode of the Data Engineering Podcast Derek Collison, creator of NATS and CEO of Synadia, talks about the evolution and capabilities of NATS as a multi-paradigm connectivity layer for distributed applications. Derek discusses the challenges and solutions in building distributed systems, and highlights the unique features of NATS that differentiate it from other messaging systems. He delves into the architectural decisions behind NATS, including its ability to handle high-speed global microservices, support for edge computing, and integration with Jetstream for data persistence, and explores the role of NATS in modern data management and its use cases in industries like manufacturing and connected vehicles.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Derek Collison about NATS, a multi-paradigm connectivity layer for distributed applications.
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what NATS is and the story behind it?
    • How have your experiences in past roles (cloud foundry, TIBCO messaging systems) informed the core principles of NATS?
      • What other sources of inspiration have you drawn on in the design and evolution of NATS? (e.g. Kafka, RabbitMQ, etc.)
    • There are several patterns and abstractions that NATS can support, many of which overlap with other well-regarded technologies. When designing a system or service, what are the heuristics that should be used to determine whether NATS should act as a replacement or addition to those capabilities? (e.g. considerations of scale, speed, ecosystem compatibility, etc.)
    • There is often a divide in the technologies and architecture used between operational/user-facing applications and data systems. How does the unification of multiple messaging patterns in NATS shift the ways that teams think about the relationship between these use cases?
      • How does the shared communication layer of NATS with multiple protocol and pattern adaptaters reduce the need to replicate data and logic across application and data layers?
    • Can you describe how the core NATS system is architected?
      • How have the design and goals of NATS evolved since you first started working on it?
    • In the time since you first began writing NATS (~2012) there have been several evolutionary stages in both application and data implementation patterns. How have those shifts influenced the direction of the NATS project and its ecosystem?
    • For teams who have an existing architecture, what are some of the patterns for adoption of NATS that allow them to augment or migrate their capabilities?
    • What are some of the ecosystem investments that you and your team have made to ease the adoption and integration of NATS?
    • What are the most interesting, innovative, or unexpected ways that you have seen NATS used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on NATS?
    • When is NATS the wrong choice?
    • What do you have planned for the future of NATS?
    Contact Info
    • GitHub
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • NATS
    • NATS JetStream
    • Synadia
    • Cloud Foundry
    • TIBCO
    • Applied Physics Lab - Johns Hopkins University
    • Cray Supercomputer
    • RVCM Certified Messaging
    • TIBCO ZMS
    • IBM MQ
    • JMS == Java Message Service
    • RabbitMQ
    • MongoDB
    • NodeJS
    • Redis
    • AMQP == Advanced Message Queueing Protocol
    • Pub/Sub Pattern
    • Circuit Breaker Pattern
    • Zero MQ
    • Akamai
    • Fastly
    • CDN == Content Delivery Network
    • At Most Once
    • At Least Once
    • Exactly Once
    • AWS Kinesis
    • Memcached
    • SQS
    • Segment
    • Rudderstack
      • Podcast Episode
    • DLQ == Dead Letter Queue
    • MQTT == Message Queueing Telemetry Transport
    • NATS Kafka Bridge
    • 10BaseT Network
    • Web Assembly
    • RedPanda
      • Podcast Episode
    • Pulsar Functions
    • mTLS
    • AuthZ (Authorization)
    • AuthN (Authentication)
    • NATS Auth Callouts
    • OPA == Open Policy Agent
    • RAG == Retrieval Augmented Generation
      • AI Engineering Podcast Episode
    • Home Assistant
      • Podcast.__init__ Episode
    • Tailscale
    • Ollama
    • CDC == Change Data Capture
    • gRPC
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 13 min
  • Advanced Lakehouse Management With The LakeKeeper Iceberg REST Catalog
    Summary
    In this episode of the Data Engineering Podcast Viktor Kessler, co-founder of Vakmo, talks about the architectural patterns in the lake house enabled by a fast and feature-rich Iceberg catalog. Viktor shares his journey from data warehouses to developing the open-source project, Lakekeeper, an Apache Iceberg REST catalog written in Rust that facilitates building lake houses with essential components like storage, compute, and catalog management. He discusses the importance of metadata in making data actionable, the evolution of data catalogs, and the challenges and innovations in the space, including integration with OpenFGA for fine-grained access control and managing data across formats and compute engines.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Viktor Kessler about architectural patterns in the lakehouse that are unlocked by a fast and feature-rich Iceberg catalog
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what LakeKeeper is and the story behind it? 
      • What is the core of the problem that you are addressing?
    • There has been a lot of activity in the catalog space recently. What are the driving forces that have highlighted the need for a better metadata catalog in the data lake/distributed data ecosystem?
      • How would you characterize the feature sets/problem spaces that different entrants are focused on addressing?
    • Iceberg as a table format has gained a lot of attention and adoption across the data ecosystem. The REST catalog format has opened the door for numerous implementations. What are the opportunities for innovation and improving user experience in that space?
    • What is the role of the catalog in managing security and governance? (AuthZ, auditing, etc.)
      • What are the channels for propagating identity and permissions to compute engines? (how do you avoid head-scratching about permission denied situations)
    • Can you describe how LakeKeeper is implemented?
      • How have the design and goals of the project changed since you first started working on it?
    • For someone who has an existing set of Iceberg tables and catalog, what does the migration process look like?
    • What new workflows or capabilities does LakeKeeper enable for data teams using Iceberg tables across one or more compute frameworks?
    • What are the most interesting, innovative, or unexpected ways that you have seen LakeKeeper used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on LakeKeeper?
    • When is LakeKeeper the wrong choice?
    • What do you have planned for the future of LakeKeeper?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • LakeKeeper
    • SAP
    • Microsoft Access
    • Microsoft Excel
    • Apache Iceberg
      • Podcast Episode
    • Iceberg REST Catalog
    • PyIceberg
    • Spark
    • Trino
    • Dremio
    • Hive Metastore
    • Hadoop
    • NATS
    • Polars
    • DuckDB
      • Podcast Episode
    • DataFusion
    • Atlan
      • Podcast Episode
    • Open Metadata
      • Podcast Episode
    • Apache Atlas
    • OpenFGA
    • Hudi
      • Podcast Episode
    • Delta Lake
      • Podcast Episode
    • Lance Table Format
      • Podcast Episode
    • Unity Catalog
    • Polaris Catalog
    • Apache Gravitino
      • Podcast Episode 
    • Keycloak
    • Open Policy Agent (OPA)
    • Apache Ranger
    • Apache NiFi
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    58 min
  • Simplifying Data Pipelines with Durable Execution
    Summary
    In this episode of the Data Engineering Podcast Jeremy Edberg, CEO of DBOS, about durable execution and its impact on designing and implementing business logic for data systems. Jeremy explains how DBOS's serverless platform and orchestrator provide local resilience and reduce operational overhead, ensuring exactly-once execution in distributed systems through the use of the Transact library. He discusses the importance of version management in long-running workflows and how DBOS simplifies system design by reducing infrastructure needs like queues and CI pipelines, making it beneficial for data pipelines, AI workloads, and agentic AI.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Jeremy Edberg about durable execution and how it influences the design and implementation of business logic
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what DBOS is and the story behind it?
    • What is durable execution?
      • What are some of the notable ways that inclusion of durable execution in an application architecture changes the ways that the rest of the application is implemented? (e.g. error handling, logic flow, etc.)
    • Many data pipelines involve complex, multi-step workflows. How does DBOS simplify the creation and management of resilient data pipelines? 
    • How does durable execution impact the operational complexity of data management systems?
    • One of the complexities in durable execution is managing code/data changes to workflows while existing executions are still processing. What are some of the useful patterns for addressing that challenge and how does DBOS help?
    • Can you describe how DBOS is architected?
      • How have the design and goals of the system changed since you first started working on it?
    • What are the characteristics of Postgres that make it suitable for the persistence mechanism of DBOS?
    • What are the guiding principles that you rely on to determine the boundaries between the open source and commercial elements of DBOS?
    • What are the most interesting, innovative, or unexpected ways that you have seen DBOS used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on DBOS?
    • When is DBOS the wrong choice?
    • What do you have planned for the future of DBOS?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • DBOS
    • Exactly Once Semantics
    • Temporal
    • Sempahore
    • Postgres
    • DBOS Transact
      • Python 
      • Typescript 
    • Idempotency Keys
    • Agentic AI
    • State Machine
    • YugabyteDB
      • Podcast Episode
    • CockroachDB
    • Supabase
    • Neon
      • Podcast Episode
    • Airflow
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    40 min
  • Overcoming Redis Limitations: The Dragonfly DB Approach
    Summary
    In this episode of the Data Engineering Podcast Roman Gershman, CTO and founder of Dragonfly DB, explores the development and impact of high-speed in-memory databases. Roman shares his experience creating a more efficient alternative to Redis, focusing on performance gains, scalability, and cost efficiency, while addressing limitations such as high throughput and low latency scenarios. He explains how Dragonfly DB solves operational complexities for users and delves into its technical aspects, including maintaining compatibility with Redis while innovating on memory efficiency. Roman discusses the importance of cost efficiency and operational simplicity in driving adoption and shares insights on the broader ecosystem of in-memory data stores, future directions like SSD tiering and vector search capabilities, and the lessons learned from building a new database engine.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Roman Gershman about building a high-speed in-memory database and the impact of the performance gains on data applications
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what DragonflyDB is and the story behind it?
    • What is the core problem/use case that is solved by making a "faster Redis"?
    • The other major player in the high performance key/value database space is Aerospike. What are the heuristics that an engineer should use to determine whether to use that vs. Dragonfly/Redis?
    • Common use cases for Redis involve application caches and queueing (e.g. Celery/RQ). What are some of the other applications that you have seen Redis/Dragonfly used for, particularly in data engineering use cases?
    • There is a piece of tribal wisdom that it takes 10 years for a database to iron out all of the kinks. At the same time, there have been substantial investments in commoditizing the underlying components of database engines. Can you describe how you approached the implementation of DragonflyDB to arive at a functional and reliable implementation?
    • What are the architectural elements that contribute to the performance and scalability benefits of Dragonfly?
      • How have the design and goals of the system changed since you first started working on it?
    • For teams who migrate from Redis to Dragonfly, beyond the cost savings what are some of the ways that it changes the ways that they think about their overall system design?
    • What are the most interesting, innovative, or unexpected ways that you have seen Dragonfly used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on DragonflyDB?
    • When is DragonflyDB the wrong choice?
    • What do you have planned for the future of DragonflyDB?
    Contact Info
    • GitHub
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • DragonflyDB
    • Redis
    • Elasticache
    • ValKey
    • Aerospike
    • Laravel
    • Sidekiq
    • Celery
    • Seastar Framework
    • Shared-Nothing Architecture
    • io_uring
    • midi-redis
    • Dunning-Kruger Effect
    • Rust
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    44 min
  • Bringing AI Into The Inner Loop of Data Engineering With Ascend
    Summary
    In this episode of the Data Engineering Podcast Sean Knapp, CEO of Ascend.io, explores the intersection of AI and data engineering. He discusses the evolution of data engineering and the role of AI in automating processes, alleviating burdens on data engineers, and enabling them to focus on complex tasks and innovation. The conversation covers the challenges and opportunities presented by AI, including the need for intelligent tooling and its potential to streamline data engineering processes. Sean and Tobias also delve into the impact of generative AI on data engineering, highlighting its ability to accelerate development, improve governance, and enhance productivity, while also noting the current limitations and future potential of AI in the field.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • Your host is Tobias Macey and today I'm interviewing Sean Knapp about how Ascend is incorporating AI into their platform to help you keep up with the rapid rate of change
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Ascend is and the story behind it?
    • The last time we spoke was August of 2022. What are the most notable or interesting evolutions in your platform since then?
      • In that same time "AI" has taken up all of the oxygen in the data ecosystem. How has that impacted the ways that you and your customers think about their priorities?
    • The introduction of AI as an API has caused many organizations to try and leap-frog their data maturity journey and jump straight to building with advanced capabilities. How is that impacting the pressures and priorities felt by data teams?
    • At the same time that AI-focused product goals are straining data teams capacities, AI also has the potential to act as an accelerator to their work. What are the roadblocks/speedbumps that are in the way of that capability?
    • Many data teams are incorporating AI tools into parts of their workflow, but it can be clunky and cumbersome. How are you thinking about the fundamental changes in how your platform works with AI at its center?
    • Can you describe the technical architecture that you have evolved toward that allows for AI to drive the experience rather than being a bolt-on?
      • What are the concrete impacts that these new capabilities have on teams who are using Ascend?
    • What are the most interesting, innovative, or unexpected ways that you have seen Ascend + AI used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on incorporating AI into the core of Ascend?
    • When is Ascend the wrong choice?
    • What do you have planned for the future of AI in Ascend?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Ascend
    • Cursor AI Code Editor
    • Devin
    • GitHub Copilot
    • OpenAI DeepResearch
    • S3 Tables
    • AWS Glue
    • AWS Bedrock
    • Snowpark
    • Co-Intelligence: Living and Working with AI by Ethan Mollick (affiliate link)
    • OpenAI o3
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    53 min

About Data Engineering Podcast

From the publisher's feed

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

More shows like Data Engineering Podcast

This Week in Startups by Jason Calacanis

This Week in Startups

1,290 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,087 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Risky Business by Risky Business Media

Risky Business

375 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

582 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

Syntax - Tasty Web Development Treats

985 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

The Data Engineering Show by The Firebolt Data Bros

The Data Engineering Show

8 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

102 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

684 Listeners