The Data Engineering Show

The Data Engineering Show

By The Firebolt Data BrosBusinessTechnologyManagement
Download on the App Store

The Data Engineering Show episodes

  • Database Technology in the Age of AI with DuckDB Labs co-creator Hannes Mühleisen
    In this episode of The Data Engineering Show, host Benjamin and co-host Eldad sit with CEO DuckDB Labs and co-creator DuckDB, Hannes Mühleisen.
    Together, they:

    • Talk about the journey of DuckDB, an open-source analytical database system designed as a universal wrangling tool.
    • Explain how DuckDB differs from SQLite, highlighting the analytical and transactional use cases.
    • Discuss DuckDB’s special feature and its approach to innovation including creating their Parquet Reader.
    • Explore the simple and efficient ecosystem of DuckDB, allowing developers to add custom functionality without changing its core stability.
    • Consider Hannes' perspective on the role of AI in databases.
    • Delve into the system’s infrastructure, design choices and the dedication of the team to ensure a continuous, reliable database system.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts, instructions on how to do this are [insert link].
    Hannes Mühleisen is the CEO of DuckDB Labs and a Professor in The Netherlands, renowned for co-creating DuckDB, an open-source analytical database system. With a background in database architecture and research from CWI database architectures group, he has pioneered the development of DuckDB as a universal data wrangling tool that can run everywhere from phones to space satellites. Under his leadership, DuckDB has achieved remarkable success, reaching 10 million downloads monthly and becoming a go-to solution for analytical database needs. His commitment to keeping DuckDB lightweight, portable, and hardware-agnostic while maintaining high performance has revolutionized how developers approach analytical database solutions. As both an academic and technology leader, Hannes brings unique insights into database architecture, open-source development, and the future of analytical data processing.
    Episode Highlights:
    • The Purpose of DuckDB (01:04)
    Hannes gives a full description of what DuckDB is as well as what it is designed to do. He describes the tool as one that understands SQL and is specifically designed to simplify complex analytical use cases.
    • SQLite vs DuckDB (02:53)
    Hannes compares two different tools stating that SQLite is an amazing system that is not meant for analytical queries but for transactional use cases while DuckDB is specifically designed for that exact purpose - analytical use cases. 
    • The Importance of Collaboration (08:14)
    Hannes states the need for community collaboration as the database engine space seems to have hundreds of brilliant people trying to solve the same problems. He shares his profound admiration for a team in Munich, praising them for their exploits in implementing concepts only described in paper.
    • The Component-Based Architecture of DuckDB (11:25)
    Hannes highlights a special feature in DuckDB, that is, it can be used as a component and he explains that the in-process architecture is a success because of the memory of data sharing that can be achieved.
    • The Parquet Reader Journey (17:51)
    Hannes explains how he built his Parquet Reader out of necessity, although he would have preferred not to. He shares how a creator named Ove Korn from Germany donated the reader to a project named “The Arrow Project” and managed it to the degree that the entire project depended on the use of the Parquet Reader and it became an issue to use both independently. Hannes adds that a parquet reader that is competent has no choice but to become a database engine which is one of the interesting things about development.
    • The Role of AI in Database Interaction (22:41)
    Hannes states that he doesn’t think that AI has a place in a database engine but rather, it is needed for optimization because the researchers who built their careers on optimization are out of jobs. He explains that the role of AI should be for assistance tasks and not for a total execution.
    • SQL - A Defined Interface (29:20)
    Hannes introduces us to a tool that allows us to pro-programmatically build a query called relational API stating that it helps to simplify the tasks of a programmer. Although, Hannes agrees that using a well-defined interface is important for components like databases, he also argues that SQL can provide a relatively defined behavior within a single system. 
    • The Golden Age of Database (38:57)
    Hannes concludes the episode by appreciating Firebolt and other engineers for taking on core engine tasks. He shares his excitement for the golden age of databases where there is a showcasing of what is possible.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.

    Quotes:
    1. “DuckDB is a universal data wrangling tool. It is a relational data management system that speaks SQL designed to do well on analytical use cases.”

    1. “We call ourselves the SQLite for analytics because it explains the original design goal of DuckDB very well.”

    1. “Within the database engine space, we are all working to solve the same problems, and that's like, a hundred of us on the planet.”

    1. “It actually turns out in order to make a competent parquet reader, you do need query execution. There is just no way around it.”

    1. “I really like this golden age of databases we are in and personally, as somebody who really likes tables and SQL, I'm quite happy to see things like firebolt and others really working on core engine stuff.”

    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository 
    • [email protected]
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    31 min
  • AI and Data Movement: Trends and Best Practices with Estuary’s Daniel Pálma
    In this episode of The Data Engineering Show, the bros sit with Daniel Pálma, Head of Marketing at Estuary.
    Join them as they:
    • Talk about Daniel’s career transition from data engineering to marketing and how his background in data engineering has been a tremendous help to his marketing competence.
    • Discuss the role of AI in the evolution of data movement ensuring a faster and easier process of creating data pipelines.
    • Shine light on the challenges of vector databases and structured data in AI applications.
    • Delve into the future of Apache Iceberg and data lakehouses, highlighting their current challenges.
    • Shares insights on the golden age of data expressing the need for more data engineers, data analysts and data practitioners in the data space.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts, instructions on how to do this are here.
    Daniel Pálma serves as Head of Marketing at Estuary, bringing a unique blend of technical expertise and marketing acumen to the data integration space. With nearly a decade of experience as a data engineer across startups, enterprises, and consulting roles, Daniel made a strategic pivot to marketing to help bridge the gap between complex technical solutions and their practical applications for data practitioners. His background in data engineering enables him to deeply understand the customers' challenges and create authentic, education-focused marketing content that resonates with technical audiences. Daniel’s thought leadership and content creation in the data engineering space, combined with his hands-on technical experience, positions him as a valuable voice in conversations about the evolution of data infrastructure and integration technologies.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository 
    • [email protected]
    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    31 min
  • AI and Data Change Management with Chad Sanderson, CEO Gable AI
    In this episode of The Data Engineering Show, host Benjamin and co-host Eldad sit with Chad Sanderson, CEO and co-founder of Gable AI  to explore the interesting world of data change management.
    Join them as they:
    • Delve into challenges of data quality, how it degrades over time and the one-sided data quality checks on the “last mile” of the data supply chain.
    • Talk about how Gable works through a 3-layer flow of technology which is to identify data production points, trace the data flow and communicate the impact of changes before they reach production.
    • Explain why the gap between data producers and consumers need to be bridged and how Gable continues to emphasize the need for effective communication and understanding data change management across teams
    • Shine light on how AI can enhance data management by extracting semantics from code and effectively manage the translation output.
    • Discuss Chad’s vision for 2025 which is to help companies start to care about data and how the changes made to data affect other people.

    Chad Sanderson is the CEO and co-founder of Gable AI, a data change management platform. Chad has over a decade of experience in data engineering and infrastructure space, holding significant roles at major companies like Microsoft, Oracle, Sephora where he focused on data quality and governance challenges. He is a former Head of Data at Convoy, a LinkedIn writer, and a published author. He lives in Seattle, Washington, and is the Chief Operator of the Data Quality Camp. His journey from data scientist to data engineer and ultimately to CEO was driven by a desire to transform how organizations manage and utilize data. Gable AI addresses the complexities of the data supply chain, by providing tools for code scanning, data contracts and governance as code, enabling teams to proactively manage data changes and impact.

    If you enjoyed this episode, make sure to subscribe, rate, and review it on Apple Podcasts, Spotify, and YouTube Podcasts. Instructions on how to do this are here.
    Episode Resources
    • Gable AI website
    • Chad Sanderson on LinkedIn

    For Feedback & Discussions on Firebolt Core:

    • Join Firebolt Discord Community
    • Join Firebolt GitHub Discussions
    • Firebolt Core Github Repository 
    • [email protected]


    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    37 min
  • Tech Stacks and Tradeoffs: Xudo's Founder on Picking the Right Tools for BI Success

    Wouter Trappers is the founder of Xudo and shares his slightly unconventional path from philosopher to data consultant with the Bros in this latest episode of The Data Engineering Show. Wouter’s grounding in philosophy has proved to be a shaping influence on his approach to business intelligence. Much more than just a software solution, for Wouter, BI is all about change management and aligning leadership with data projects.


    They discuss:


    • From Excel to Expert:
      From basic Excel tasks to a full mastery of BI tools like QlikView, Wouter has blended his technical and philosophical approaches to data to become a bona fide expert.
    • Data Strategy as Transformation: Good change management principles have to be adhered to if a BI project is going to bear fruit. Focus on leadership alignment, KPI clarity, and user empowerment instead of simply implementing software. 
    • Challenges of Starting Small: Wouter has some tips to offer smaller companies around bootstrapping their data journey using existing tools, practical education, and even Gen AI.
    • Balancing Scales: Smaller startups compared to large enterprises face a very different set of challenges.


    Wouter’s combination of philosophy and pragmatism brings fresh takes to building effective data solutions.



    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    25 min
  • Data Rewind: Conversation Highlights from Zach Wilson, Matthew Housley, Joe Reis, and Krishnan Viswanathan

    In this special roundup episode of The Data Engineering Show, the Bros revisits some of the best bits from episodes with data thought leaders Zach Wilson, Matthew Housley, Joe Reis, and Krishnan Viswanathan, spotlighting essential trends and lessons learned across the evolving data engineering landscape. From data observability to bridging academia with real-world practice, this episode covers perspectives on where data engineering is heading and why certain challenges persist.


    Topics include:

    • Foundations of Data Engineering: Zach Wilson emphasizes the importance of core, tech-agnostic skills in data modeling, quality assurance, and storytelling. By sharing his experiences at Airbnb and in education, he reveals that effective data engineering hinges on creating robust data models, quality controls, and persuasive narratives rather than expertise in any single tool or language.
    • Bridging Academia and Practice: Matthew Housley and Joe Reis delve into the need for better data education, emphasizing hands-on experience and data fundamentals over tool-specific training, and advocate for apprenticeships and real-world collaborations in educational settings.
    • Legacy Meets Modern in Data Engineering: Krishnan Viswanathan reflects on recurring themes in data engineering and the importance of adapting legacy approaches to new data needs, underscoring the challenges and benefits of vendor-built versus in-house solutions.


    Join the Bros for a well-rounded exploration of current themes in data engineering, filled with practical advice for data professionals at any stage of their journey.



    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    29 min
  • The Resurgence of SQL: Insights from Ryanne Dolan from LinkedIn

    In this episode of The Data Engineering Show, the bros, Eldad and Benjamin are joined by Ryanne Dolan from LinkedIn to discuss the innovative Hoptimator (H2) project. This conversation reveals how LinkedIn has improved its data pipelines by automating the setup and management of complex workflows.


    Together they cover:

    • Automated Data Pipelines: Ryanne explains how Hoptimator allows users to create and manage data pipelines using just a simple SQL SELECT query, streamlining the process of setting up Kafka topics, Flink jobs, and schemas.
    • Integration with Kubernetes: The project utilizes Kubernetes to handle infrastructure tasks, treating Kubernetes as a database for managing state. This integration simplifies the orchestration of data workflows and automates routine tasks.
    • Consumer-Driven Model: Ryanne discusses the shift from a producer-driven to a consumer-driven data model, emphasizing the importance of understanding and addressing consumer needs to reduce engineering complexity and optimize data systems.
    • Future of Data Engineering: The conversation touches on the ongoing experimental nature of Hoptimator and its potential to transform data engineering practices, highlighting its impact on LinkedIn's data infrastructure.


    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    33 min
  • Vector Databases Won’t Replace SQL - Andy Pavlo

    SQL’s slow. SQL’s stupid. We hear these claims every time a new shiny tool enters the market, only to realize five years later when the hype dies down that SQL is actually a good idea. 

    In this super techie episode of the Data Engineering Show, Andy Pavlo, Associate Professor at Carnegie Mellon University, joins the bros to delve into database internals and optimization. 

    Andy discusses leveraging ML for autonomous database optimization, using Postgres for practical applications, tuning production databases safely, and why SQL is here to stay.

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    43 min
  • How ZoomInfo transitioned from data graveyards to ROI-driven data projects

    Too often expensive resources and manhours are spent on dashboards no one uses, resulting in zero ROI. Philip Philip Zelitchenko, VP of Data & Analytics at ZoomInfo met the bros to talk about adopting product management principles to ensure data projects have value, and provide an unfiltered peak into ZoomInfo’s data stack and unique tech culture. 

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    40 min
  • Matthew Weingarten from Disney Streaming about Data Quality Best Practices

    Matthew Weingarten, Lead Data Engineer at Disney Streaming, talks about principles essential for data quality, cost optimization, debugging, and data modeling, as adopted by the world's leading companies.

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    28 min
  • Joseph Machado, Senior Data Engineer @ LinkedIn talks best practices

    Data engineering should be less about the stack and more about best practices. While tools may change, foundational principles will remain constant. Joseph Mercado, Senior Data Engineer at LinkedIn, is on The Data Engineering Show to talk about principles that are key to success, leveraging AI for automation, and adopting software engineering methods. 

    The Data Engineering Show is brought to you by firebolt.io and handcrafted by our friends over at: fame.so

    Previous guests include: Joseph Machado of Linkedin, Metthew Weingarten of Disney, Joe Reis and Matt Housely, authors of The Fundamentals of Data Engineering, Zach Wilson of Eczachly Inc, Megan Lieu of Deepnote, Erik Heintare of Bolt, Lior Solomon of Vimeo, Krishna Naidu of Canva, Mike Cohen of Substack, Jens Larsson of Ark, Gunnar Tangring of Klarna, Yoav Shmaria of Similarweb and Xiaoxu Gao of Adyen.

    Check out our three most downloaded episodes:
    • Zach Wilson on What Makes a Great Data Engineer
    • Joe Reis and Matt Housley on The Fundamentals of Data Engineering
    • Bill Inmon, The Godfather of Data Warehousing
    26 min

About The Data Engineering Show

From the publisher's feed

The Data Engineering Show is a podcast for data engineering and BI practitioners to go beyond theory. Learn from the biggest influencers in tech about their practical day-to-day data challenges and solutions in a casual and fun setting.

More shows like The Data Engineering Show

Planet Money by NPR

Planet Money

30,701 Listeners

Hidden Brain by Hidden Brain, Shankar Vedantam

Hidden Brain

43,359 Listeners

Data Engineering Podcast by Tobias Macey

Data Engineering Podcast

144 Listeners

DataFramed by DataCamp

DataFramed

265 Listeners

Tech Brew Ride Home by Morning Brew

Tech Brew Ride Home

959 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

The Journal. by The Wall Street Journal & Spotify Studios

The Journal.

6,068 Listeners

My First Million by Hubspot Media

My First Million

2,647 Listeners

The Prof G Pod with Scott Galloway by Vox Media Podcast Network

The Prof G Pod with Scott Galloway

5,391 Listeners

The Real Python Podcast by Real Python

The Real Python Podcast

139 Listeners

All-In with Chamath, Jason, Sacks & Friedberg by All-In Podcast, LLC

All-In with Chamath, Jason, Sacks & Friedberg

10,186 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

The Analytics Engineering Podcast by dbt Labs, Inc.

The Analytics Engineering Podcast

29 Listeners

HBR On Leadership by Harvard Business Review

HBR On Leadership

156 Listeners

Training Data by Sequoia Capital

Training Data

39 Listeners