May 15, 2023

What Happens When The Abstractions Leak On Your Data

26 minutes

Summary

All of the advancements in our technology is based around the principles of abstraction. These are valuable until they break down, which is an inevitable occurrence. In this episode the host Tobias Macey shares his reflections on recent experiences where the abstractions leaked and some observances on how to deal with that situation in a data platform architecture.

Announcements

Hello and welcome to the Data Engineering Podcast, the show about modern data management

RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control. Their SDKs make event streaming from any app or website easy, and their extensive library of integrations enable you to automatically send data to hundreds of downstream tools. Sign up free at dataengineeringpodcast.com/rudderstack

Your host is Tobias Macey and today I'm sharing some thoughts and observances about abstractions and impedance mismatches from my experience building a data lakehouse with an ELT workflow

Interview

Introduction

impact of community tech debt

hive metastore

new work being done but not widely adopted

tensions between automation and correctness

data type mapping

integer types

complex types

naming things (keys/column names from APIs to databases)

disaggregated databases - pros and cons

flexibility and cost control

not as much tooling invested vs. Snowflake/BigQuery/Redshift

data modeling

dimensional modeling vs. answering today's questions

What are the most interesting, unexpected, or challenging lessons that you have learned while working on your data platform?

When is ELT the wrong choice?

What do you have planned for the future of your data platform?

Contact Info

Parting Question

From your perspective, what is the biggest gap in the tooling or technology for data management today?

Closing Announcements

Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.

Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.

If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.

To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers

Links

dbt

Airbyte

Podcast Episode

Dagster

Podcast Episode

Trino

Podcast Episode

ELT

Data Lakehouse

Snowflake

BigQuery

Redshift

Technical Debt

Hive Metastore

AWS Glue

The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

What Happens When The Abstractions Leak On Your Data

26 minutes

Summary

Announcements

Hello and welcome to the Data Engineering Podcast, the show about modern data management

Your host is Tobias Macey and today I'm sharing some thoughts and observances about abstractions and impedance mismatches from my experience building a data lakehouse with an ELT workflow

Interview

Introduction

impact of community tech debt

hive metastore

new work being done but not widely adopted

tensions between automation and correctness

data type mapping

integer types

complex types

naming things (keys/column names from APIs to databases)

disaggregated databases - pros and cons

flexibility and cost control

not as much tooling invested vs. Snowflake/BigQuery/Redshift

data modeling

dimensional modeling vs. answering today's questions

What are the most interesting, unexpected, or challenging lessons that you have learned while working on your data platform?

When is ELT the wrong choice?

What do you have planned for the future of your data platform?

Contact Info

Parting Question

From your perspective, what is the biggest gap in the tooling or technology for data management today?

Closing Announcements

Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.

If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.

To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers

Links

dbt

Airbyte

Podcast Episode

Dagster

Podcast Episode

Trino

Podcast Episode

ELT

Data Lakehouse

Snowflake

BigQuery

Redshift

Technical Debt

Hive Metastore

AWS Glue

The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

More shows like Data Engineering Podcast

View all

This Week in Startups

1,290 Listeners

The Changelog: Software Development, Open Source

289 Listeners

The a16z Show

1,093 Listeners

Software Engineering Daily

626 Listeners

Risky Business

375 Listeners

Talk Python To Me

583 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn

301 Listeners

NVIDIA AI Podcast

345 Listeners

Syntax - Tasty Web Development Treats

982 Listeners

Practical AI

208 Listeners

Dwarkesh Podcast

576 Listeners

The Data Engineering Show

8 Listeners

Latent Space: The AI Engineer Podcast

101 Listeners

This Day in AI Podcast

226 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis

682 Listeners

Share What Happens When The Abstractions Leak On Your Data

Sign up to save your podcasts

What Happens When The Abstractions Leak On Your Data

What Happens When The Abstractions Leak On Your Data

More shows like Data Engineering Podcast

This Week in Startups

The Changelog: Software Development, Open Source

The a16z Show

Software Engineering Daily

Risky Business

Talk Python To Me

Super Data Science: ML & AI Podcast with Jon Krohn

NVIDIA AI Podcast

Syntax - Tasty Web Development Treats

Practical AI

Dwarkesh Podcast

The Data Engineering Show

Latent Space: The AI Engineer Podcast

This Day in AI Podcast

The AI Daily Brief: Artificial Intelligence News and Analysis