Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Unpacking The Seven Principles Of Modern Data Pipelines
    Summary

    Data pipelines are the core of every data product, ML model, and business intelligence dashboard. If you're not careful you will end up spending all of your time on maintenance and fire-fighting. The folks at Rivery distilled the seven principles of modern data pipelines that will help you stay out of trouble and be productive with your data. In this episode Ariel Pohoryles explains what they are and how they work together to increase your chances of success.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
    • This episode is brought to you by Datafold – a testing automation platform for data engineers that finds data quality issues before the code and data are deployed to production. Datafold leverages data-diffing to compare production and development environments and column-level lineage to show you the exact impact of every code change on data, metrics, and BI tools, keeping your team productive and stakeholders happy. Datafold integrates with dbt, the modern data stack, and seamlessly plugs in your data CI for team-wide and automated testing. If you are migrating to a modern data stack, Datafold can also help you automate data and code validation to speed up the migration. Learn more about Datafold by visiting dataengineeringpodcast.com/datafold
    • Your host is Tobias Macey and today I'm interviewing Ariel Pohoryles about the seven principles of modern data pipelines
    • Interview
      • Introduction
      • How did you get involved in the area of data management?
      • Can you start by defining what you mean by a "modern" data pipeline?
      • At Rivery you published a white paper identifying seven principles of modern data pipelines:
        • Zero infrastructure management
        • ELT-first mindset
        • Speaks SQL and Python
        • Dynamic multi-storage layers
        • Reverse ETL & operational analytics
        • Full transparency
        • Faster time to value
        • What are the applications of data that you focused on while identifying these principles?
        • How do the application of these principles influence the ability of organizations and their data teams to encourage and keep pace with the use of data in the business?
        • What are the technical components of a pipeline infrastructure that are necessary to support a "modern" workflow?
        • How do the technologies involved impact the organizational involvement with how data is applied throughout the business?
        • When using managed services, what are the ways that the pricing model acts to encourage/discourage experimentation/exploration with data?
        • What are the most interesting, innovative, or unexpected ways that you have seen these seven principles implemented/applied?
        • What are the most interesting, unexpected, or challenging lessons that you have learned while working with customers to adapt to these principles?
        • What are the cases where some/all of these principles are undesirable/impractical to implement?
        • What are the opportunities for further advancement/sophistication in the ways that teams work with and gain value from data?
        • Contact Info
          • LinkedIn
          • Parting Question
            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
            • Closing Announcements
              • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
              • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
              • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
              • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
              • Links
                • Rivery
                • 7 Principles Of The Modern Data Pipeline
                • ELT
                • Reverse ETL
                • Martech Landscape
                • Data Lakehouse
                • Databricks
                • Snowflake
                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                  Sponsored By:

                  • Datafold: ![Datafold](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zm6x2tFu.png)
                  This episode is brought to you by Datafold – a testing automation platform for data engineers that finds data quality issues before the code and data are deployed to production. Datafold leverages data-diffing to compare production and development environments and column-level lineage to show you the exact impact of every code change on data, metrics, and BI tools, keeping your team productive and stakeholders happy. Datafold integrates with dbt, the modern data stack, and seamlessly plugs in your data CI for team-wide and automated testing. If you are migrating to a modern data stack, Datafold can also help you automate data and code validation to speed up the migration. Learn more about Datafold by visiting [dataengineeringpodcast.com/datafold](https://www.dataengineeringpodcast.com/datafold) today!
                • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                  Support Data Engineering Podcast

                  48 min
                • Quantifying The Return On Investment For Your Data Team
                  Summary

                  As businesses increasingly invest in technology and talent focused on data engineering and analytics, they want to know whether they are benefiting. So how do you calculate the return on investment for data? In this episode Barr Moses and Anna Filippova explore that question and provide useful exercises to start answering that in your company.

                  Announcements
                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                  • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                  • Your host is Tobias Macey and today I'm interviewing Barr Moses and Anna Filippova about how and whether to measure the ROI of your data team
                  • Interview
                    • Introduction
                    • How did you get involved in the area of data management?
                    • What are the typical motivations for measuring and tracking the ROI for a data team?
                      • Who is responsible for collecting that information?
                      • How is that information used and by whom?
                      • What are some of the downsides/risks of tracking this metric? (law of unintended consequences)
                      • What are the inputs to the number that constitutes the "investment"? infrastructure, payroll of employees on team, time spent working with other teams?
                      • What are the aspects of data work and its impact on the business that complicate a calculation of the "return" that is generated?
                      • How should teams think about measuring data team ROI?
                      • What are some concrete ROI metrics data teams can use?
                        • What level of detail is useful? What dimensions should be used for segmenting the calculations?
                        • How can visibility into this ROI metric be best used to inform the priorities and project scopes of the team?
                        • With so many tools in the modern data stack today, what is the role of technology in helping drive or measure this impact?
                        • How do your respective solutions, Monte Carlo and dbt, help teams measure and scale data value?
                        • With generative AI on the upswing of the hype cycle, what are the impacts that you see it having on data teams?
                          • What are the unrealistic expectations that it will produce?
                          • How can it speed up time to delivery?
                          • What are the most interesting, innovative, or unexpected ways that you have seen data team ROI calculated and/or used?
                          • What are the most interesting, unexpected, or challenging lessons that you have learned while working on measuring the ROI of data teams?
                          • When is measuring ROI the wrong choice?
                          • Contact Info
                            • Barr
                              • LinkedIn
                              • Anna
                                • LinkedIn
                                • Parting Question
                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                  • Closing Announcements
                                    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                    • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                    • Links
                                      • Monte Carlo
                                        • Podcast Episode
                                        • dbt
                                          • Podcast Episode
                                          • JetBlue Snowflake Con Presentation
                                          • Generative AI
                                          • Large Language Models
                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                            Sponsored By:

                                            • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                            Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                            Support Data Engineering Podcast

                                            1 hr 2 min
                                          • Strategies For A Successful Data Platform Migration
                                            Summary

                                            All software systems are in a constant state of evolution. This makes it impossible to select a truly future-proof technology stack for your data platform, making an eventual migration inevitable. In this episode Gleb Mezhanskiy and Rob Goretsky share their experiences leading various data platform migrations, and the hard-won lessons that they learned so that you don't have to.

                                            Announcements
                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                            • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                                            • Modern data teams are using Hex to 10x their data impact. Hex combines a notebook style UI with an interactive report builder. This allows data teams to both dive deep to find insights and then share their work in an easy-to-read format to the whole org. In Hex you can use SQL, Python, R, and no-code visualization together to explore, transform, and model data. Hex also has AI built directly into the workflow to help you generate, edit, explain and document your code. The best data teams in the world such as the ones at Notion, AngelList, and Anthropic use Hex for ad hoc investigations, creating machine learning models, and building operational dashboards for the rest of their company. Hex makes it easy for data analysts and data scientists to collaborate together and produce work that has an impact. Make your data team unstoppable with Hex. Sign up today at dataengineeringpodcast.com/hex to get a 30-day free trial for your team!
                                            • Your host is Tobias Macey and today I'm interviewing Gleb Mezhanskiy and Rob Goretsky about when and how to think about migrating your data stack
                                            • Interview
                                              • Introduction
                                              • How did you get involved in the area of data management?
                                              • A migration can be anything from a minor task to a major undertaking. Can you start by describing what constitutes a migration for the purposes of this conversation?
                                              • Is it possible to completely avoid having to invest in a migration?
                                              • What are the signals that point to the need for a migration?
                                                • What are some of the sources of cost that need to be accounted for when considering a migration? (both in terms of doing one, and the costs of not doing one)
                                                • What are some signals that a migration is not the right solution for a perceived problem?
                                                • Once the decision has been made that a migration is necessary, what are the questions that the team should be asking to determine the technologies to move to and the sequencing of execution?
                                                • What are the preceding tasks that should be completed before starting the migration to ensure there is no breakage downstream of the changing component(s)?
                                                • What are some of the ways that a migration effort might fail?
                                                • What are the major pitfalls that teams need to be aware of as they work through a data platform migration?
                                                • What are the opportunities for automation during the migration process?
                                                • What are the most interesting, innovative, or unexpected ways that you have seen teams approach a platform migration?
                                                • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data platform migrations?
                                                • What are some ways that the technologies and patterns that we use can be evolved to reduce the cost/impact/need for migraitons?
                                                • Contact Info
                                                  • Gleb
                                                    • LinkedIn
                                                    • @glebmm on Twitter
                                                    • Rob
                                                      • LinkedIn
                                                      • RobGoretsky on GitHub
                                                      • Parting Question
                                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                        • Closing Announcements
                                                          • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                          • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                          • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                          • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                          • Links
                                                            • Datafold
                                                              • Podcast Episode
                                                              • Informatica
                                                              • Airflow
                                                              • Snowflake
                                                                • Podcast Episode
                                                                • Redshift
                                                                • Eventbrite
                                                                • Teradata
                                                                • BigQuery
                                                                • Trino
                                                                • EMR == Elastic Map-Reduce
                                                                • Shadow IT
                                                                  • Podcast Episode
                                                                  • Mode Analytics
                                                                  • Looker
                                                                  • Sunk Cost Fallacy
                                                                  • data-diff
                                                                    • Podcast Episode
                                                                    • SQLGlot
                                                                    • [Dagster](dhttps://dagster.io/)
                                                                    • dbt
                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                      Sponsored By:

                                                                      • Hex: ![Hex Tech Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zBEUGheK.png)
                                                                      Hex is a collaborative workspace for data science and analytics. A single place for teams to explore, transform, and visualize data into beautiful interactive reports. Use SQL, Python, R, no-code and AI to find and share insights across your organization. Empower everyone in an organization to make an impact with data. Sign up today at [dataengineeringpodcast.com/hex](https://www.dataengineeringpodcast.com/hex} and get 30 days free!
                                                                    • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                    • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                                                      Support Data Engineering Podcast

                                                                      1 hr 10 min
                                                                    • Build Real Time Applications With Operational Simplicity Using Dozer
                                                                      Summary

                                                                      Real-time data processing has steadily been gaining adoption due to advances in the accessibility of the technologies involved. Despite that, it is still a complex set of capabilities. To bring streaming data in reach of application engineers Matteo Pelati helped to create Dozer. In this episode he explains how investing in high performance and operationally simplified streaming with a familiar API can yield significant benefits for software and data teams together.

                                                                      Announcements
                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                      • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                                                                      • Modern data teams are using Hex to 10x their data impact. Hex combines a notebook style UI with an interactive report builder. This allows data teams to both dive deep to find insights and then share their work in an easy-to-read format to the whole org. In Hex you can use SQL, Python, R, and no-code visualization together to explore, transform, and model data. Hex also has AI built directly into the workflow to help you generate, edit, explain and document your code. The best data teams in the world such as the ones at Notion, AngelList, and Anthropic use Hex for ad hoc investigations, creating machine learning models, and building operational dashboards for the rest of their company. Hex makes it easy for data analysts and data scientists to collaborate together and produce work that has an impact. Make your data team unstoppable with Hex. Sign up today at dataengineeringpodcast.com/hex to get a 30-day free trial for your team!
                                                                      • Your host is Tobias Macey and today I'm interviewing Matteo Pelati about Dozer, an open source engine that includes data ingestion, transformation, and API generation for real-time sources
                                                                      • Interview
                                                                        • Introduction
                                                                        • How did you get involved in the area of data management?
                                                                        • Can you describe what Dozer is and the story behind it?
                                                                          • What was your decision process for building Dozer as open source?
                                                                          • As you note in the documentation, Dozer has overlap with a number of technologies that are aimed at different use cases. What was missing from each of them and the center of their Venn diagram that prompted you to build Dozer?
                                                                          • In addition to working in an interesting technological cross-section, you are also targeting a disparate group of personas. Who are you building Dozer for and what were the motivations for that vision?
                                                                            • What are the different use cases that you are focused on supporting?
                                                                            • What are the features of Dozer that enable engineers to address those uses, and what makes it preferable to existing alternative approaches?
                                                                            • Can you describe how Dozer is implemented?
                                                                              • How have the design and goals of the platform changed since you first started working on it?
                                                                              • What are the architectural "-ilities" that you are trying to optimize for?
                                                                              • What is involved in getting Dozer deployed and integrated into an existing application/data infrastructure?
                                                                              • How can teams who are using Dozer extend/integrate with Dozer?
                                                                                • What does the development/deployment workflow look like for teams who are building on top of Dozer?
                                                                                • What is your governance model for Dozer and balancing the open source project against your business goals?
                                                                                • What are the most interesting, innovative, or unexpected ways that you have seen Dozer used?
                                                                                • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Dozer?
                                                                                • When is Dozer the wrong choice?
                                                                                • What do you have planned for the future of Dozer?
                                                                                • Contact Info
                                                                                  • LinkedIn
                                                                                  • @pelatimtt on Twitter
                                                                                  • Parting Question
                                                                                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                    • Closing Announcements
                                                                                      • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                      • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                      • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                      • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                      • Links
                                                                                        • Dozer
                                                                                        • Data Robot
                                                                                        • Netflix Bulldozer
                                                                                        • CubeJS
                                                                                          • Podcast Episode
                                                                                          • JVM == Java Virtual Machine
                                                                                          • Flink
                                                                                            • Podcast Episode
                                                                                            • Airbyte
                                                                                              • Podcast Episode
                                                                                              • Fivetran
                                                                                                • Podcast Episode
                                                                                                • Delta Lake
                                                                                                  • Podcast Episode
                                                                                                  • LMDB
                                                                                                  • Vector Database
                                                                                                  • LLM == Large Language Model
                                                                                                  • Rockset
                                                                                                    • Podcast Episode
                                                                                                    • Tinybird
                                                                                                      • Podcast Episode
                                                                                                      • Rust Language
                                                                                                      • Materialize
                                                                                                        • Podcast Episode
                                                                                                        • RisingWave
                                                                                                        • DuckDB
                                                                                                          • Podcast Episode
                                                                                                          • DataFusion
                                                                                                          • Polars
                                                                                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                            Sponsored By:

                                                                                                            • Hex: ![Hex Tech Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zBEUGheK.png)
                                                                                                            Hex is a collaborative workspace for data science and analytics. A single place for teams to explore, transform, and visualize data into beautiful interactive reports. Use SQL, Python, R, no-code and AI to find and share insights across your organization. Empower everyone in an organization to make an impact with data. Sign up today at [dataengineeringpodcast.com/hex](https://www.dataengineeringpodcast.com/hex} and get 30 days free!
                                                                                                          • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                          • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                                                                                            Support Data Engineering Podcast

                                                                                                            41 min
                                                                                                          • Datapreneurs - How Todays Business Leaders Are Using Data To Define The Future
                                                                                                            Summary

                                                                                                            Data has been one of the most substantial drivers of business and economic value for the past few decades. Bob Muglia has had a front-row seat to many of the major shifts driven by technology over his career. In his recent book "Datapreneurs" he reflects on the people and businesses that he has known and worked with and how they relied on data to deliver valuable services and drive meaningful change.

                                                                                                            Announcements
                                                                                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                            • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                                                                                                            • Your host is Tobias Macey and today I'm interviewing Bob Muglia about his recent book about the idea of "Datapreneurs" and the role of data in the modern economy
                                                                                                            • Interview
                                                                                                              • Introduction
                                                                                                              • How did you get involved in the area of data management?
                                                                                                              • Can you describe what your concept of a "Datapreneur" is?
                                                                                                                • How is this distinct from the common idea of an entreprenur?
                                                                                                                • What do you see as the key inflection points in data technologies and their impacts on business capabilities over the past ~30 years?
                                                                                                                • In your role as the CEO of Snowflake you had a first-row seat for the rise of the "modern data stack". What do you see as the main positive and negative impacts of that paradigm?
                                                                                                                  • What are the key issues that are yet to be solved in that ecosmnjjystem?
                                                                                                                  • For technologists who are thinking about launching new ventures, what are the key pieces of advice that you would like to share?
                                                                                                                  • What do you see as the short/medium/long-term impact of AI on the technical, business, and societal arenas?
                                                                                                                  • What are the most interesting, innovative, or unexpected ways that you have seen business leaders use data to drive their vision?
                                                                                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on the Datapreneurs book?
                                                                                                                  • What are your key predictions for the future impact of data on the technical/economic/business landscapes?
                                                                                                                  • Contact Info
                                                                                                                    • LinkedIn
                                                                                                                    • Parting Question
                                                                                                                      • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                      • Closing Announcements
                                                                                                                        • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                        • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                        • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                        • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                        • Links
                                                                                                                          • Datapreneurs Book
                                                                                                                          • SQL Server
                                                                                                                          • Snowflake
                                                                                                                          • Z80 Processor
                                                                                                                          • Navigational Database
                                                                                                                          • System R
                                                                                                                          • Redshift
                                                                                                                          • Microsoft Fabric
                                                                                                                          • Databricks
                                                                                                                          • Looker
                                                                                                                          • Fivetran
                                                                                                                            • Podcast Episode
                                                                                                                            • Databricks Unity Catalog
                                                                                                                            • RelationalAI
                                                                                                                            • 6th Normal Form
                                                                                                                            • Pinecone Vector DB
                                                                                                                              • Podcast Episode
                                                                                                                              • Perplexity AI
                                                                                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                Sponsored By:

                                                                                                                                • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                                                                                                                Support Data Engineering Podcast

                                                                                                                                55 min
                                                                                                                              • Reduce Friction In Your Business Analytics Through Entity Centric Data Modeling
                                                                                                                                Summary

                                                                                                                                For business analytics the way that you model the data in your warehouse has a lasting impact on what types of questions can be answered quickly and easily. The major strategies in use today were created decades ago when the software and hardware for warehouse databases were far more constrained. In this episode Maxime Beauchemin of Airflow and Superset fame shares his vision for the entity-centric data model and how you can incorporate it into your own warehouse design.

                                                                                                                                Announcements
                                                                                                                                • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                                                                                                                                • Your host is Tobias Macey and today I'm interviewing Max Beauchemin about the concept of entity-centric data modeling for analytical use cases
                                                                                                                                • Interview
                                                                                                                                  • Introduction
                                                                                                                                  • How did you get involved in the area of data management?
                                                                                                                                  • Can you describe what entity-centric modeling (ECM) is and the story behind it?

                                                                                                                                    • How does it compare to dimensional modeling strategies?
                                                                                                                                    • What are some of the other competing methods
                                                                                                                                    • Comparison to activity schema
                                                                                                                                    • What impact does this have on ML teams? (e.g. feature engineering)

                                                                                                                                    • What role does the tooling of a team have in the ways that they end up thinking about modeling? (e.g. dbt vs. informatica vs. ETL scripts, etc.)

                                                                                                                                      • What is the impact on the underlying compute engine on the modeling strategies used?
                                                                                                                                      • What are some examples of data sources or problem domains for which this approach is well suited?

                                                                                                                                        • What are some cases where entity centric modeling techniques might be counterproductive?
                                                                                                                                        • What are the ways that the benefits of ECM manifest in use cases that are down-stream from the warehouse?

                                                                                                                                        • What are some concrete tactical steps that teams should be thinking about to implement a workable domain model using entity-centric principles?

                                                                                                                                          • How does this work across business domains within a given organization (especially at "enterprise" scale)?
                                                                                                                                          • What are the most interesting, innovative, or unexpected ways that you have seen ECM used?

                                                                                                                                          • What are the most interesting, unexpected, or challenging lessons that you have learned while working on ECM?

                                                                                                                                          • When is ECM the wrong choice?

                                                                                                                                          • What are your predictions for the future direction/adoption of ECM or other modeling techniques?

                                                                                                                                          • Contact Info
                                                                                                                                            • mistercrunch on GitHub
                                                                                                                                            • LinkedIn
                                                                                                                                            • Parting Question
                                                                                                                                              • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                              • Closing Announcements
                                                                                                                                                • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                                                • Links
                                                                                                                                                  • Entity Centric Modeling Blog Post
                                                                                                                                                  • Max's Previous Apperances
                                                                                                                                                    • Defining Data Engineering with Maxime Beauchemin
                                                                                                                                                    • Self Service Data Exploration And Dashboarding With Superset
                                                                                                                                                    • Exploring The Evolving Role Of Data Engineers
                                                                                                                                                    • Alumni Of AirBnB's Early Years Reflect On What They Learned About Building Data Driven Organizations
                                                                                                                                                    • Apache Airflow
                                                                                                                                                    • Apache Superset
                                                                                                                                                    • Preset
                                                                                                                                                    • Ubisoft
                                                                                                                                                    • Ralph Kimball
                                                                                                                                                    • The Rise Of The Data Engineer
                                                                                                                                                    • The Downfall Of The Data Engineer
                                                                                                                                                    • The Rise Of The Data Scientist
                                                                                                                                                    • Dimensional Data Modeling
                                                                                                                                                    • Star Schema
                                                                                                                                                    • Database Normalization
                                                                                                                                                    • Feature Engineering
                                                                                                                                                    • DRY == Don't Repeat Yourself
                                                                                                                                                    • Activity Schema
                                                                                                                                                      • Podcast Episode
                                                                                                                                                      • Corporate Information Factory (affiliate link)
                                                                                                                                                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                        Sponsored By:

                                                                                                                                                        • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                                        Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                                                                                                                                        Support Data Engineering Podcast

                                                                                                                                                        1 hr 13 min
                                                                                                                                                      • How Data Engineering Teams Power Machine Learning With Feature Platforms
                                                                                                                                                        Summary

                                                                                                                                                        Feature engineering is a crucial aspect of the machine learning workflow. To make that possible, there are a number of technical and procedural capabilities that must be in place first. In this episode Razi Raziuddin shares how data engineering teams can support the machine learning workflow through the development and support of systems that empower data scientists and ML engineers to build and maintain their own features.

                                                                                                                                                        Announcements
                                                                                                                                                        • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                        • Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at dataengineeringpodcast.com/rudderstack
                                                                                                                                                        • Your host is Tobias Macey and today I'm interviewing Razi Raziuddin about how data engineers can empower data scientists to develop and deploy better ML models through feature engineering
                                                                                                                                                        • Interview
                                                                                                                                                          • Introduction
                                                                                                                                                          • How did you get involved in the area of data management?
                                                                                                                                                          • What is feature engineering is and why/to whom it matters?
                                                                                                                                                            • A topic that commonly comes up in relation to feature engineering is the importance of a feature store. What are the tradeoffs for that to be a separate infrastructure/architecture component?
                                                                                                                                                            • What is the overall lifecycle of a feature, from definition to deployment and maintenance?
                                                                                                                                                              • How is this distinct from other forms of data pipeline development and delivery?
                                                                                                                                                              • Who are the participants in that workflow?
                                                                                                                                                              • What are the sharp edges/roadblocks that typically manifest in that lifecycle?
                                                                                                                                                              • What are the interfaces that are needed for data scientists/ML engineers to be able to self-serve their feature management?
                                                                                                                                                                • What is the role of the data engineer in supporting those interfaces?
                                                                                                                                                                • What are the communication/collaboration channels that are necessary to make the overall process a success?
                                                                                                                                                                • From an implementation/architecture perspective, what are the patterns that you have seen teams build around for feature development/serving?
                                                                                                                                                                • What are the most interesting, innovative, or unexpected ways that you have seen feature platforms used?
                                                                                                                                                                • What are the most interesting, unexpected, or challenging lessons that you have learned while working on feature engineering?
                                                                                                                                                                • What are the resources that you find most helpful in understanding and designing feature platforms?
                                                                                                                                                                • Contact Info
                                                                                                                                                                  • LinkedIn
                                                                                                                                                                  • Parting Question
                                                                                                                                                                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                    • Closing Announcements
                                                                                                                                                                      • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                      • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                      • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                      • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                                                                      • Links
                                                                                                                                                                        • FeatureByte
                                                                                                                                                                        • DataRobot
                                                                                                                                                                        • Feature Store
                                                                                                                                                                        • Feast Feature Store
                                                                                                                                                                        • Feathr
                                                                                                                                                                        • Kaggle
                                                                                                                                                                        • Yann LeCun
                                                                                                                                                                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                          Sponsored By:

                                                                                                                                                                          • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                                                          Introducing RudderStack Profiles. RudderStack Profiles takes the SaaS guesswork and SQL grunt work out of building complete customer profiles so you can quickly ship actionable, enriched data to every downstream team. You specify the customer traits, then Profiles runs the joins and computations for you to create complete customer profiles. Get all of the details and try the new product today at [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack)

                                                                                                                                                                          Support Data Engineering Podcast

                                                                                                                                                                          1 hr 4 min
                                                                                                                                                                        • Seamless SQL And Python Transformations For Data Engineers And Analysts With SQLMesh
                                                                                                                                                                          Summary

                                                                                                                                                                          Data transformation is a key activity for all of the organizational roles that interact with data. Because of its importance and outsized impact on what is possible for downstream data consumers it is critical that everyone is able to collaborate seamlessly. SQLMesh was designed as a unifying tool that is simple to work with but powerful enough for large-scale transformations and complex projects. In this episode Toby Mao explains how it works, the importance of automatic column-level lineage tracking, and how you can start using it today.

                                                                                                                                                                          Announcements
                                                                                                                                                                          • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                                          • RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control. Their SDKs make event streaming from any app or website easy, and their extensive library of integrations enable you to automatically send data to hundreds of downstream tools. Sign up free at dataengineeringpodcast.com/rudderstack-
                                                                                                                                                                          • Your host is Tobias Macey and today I'm interviewing Toby Mao about SQLMesh, an open source DataOps framework designed to scale data transformations with ease of collaboration and validation built in
                                                                                                                                                                          • Interview
                                                                                                                                                                            • Introduction
                                                                                                                                                                            • How did you get involved in the area of data management?
                                                                                                                                                                            • Can you describe what SQLMesh is and the story behind it?
                                                                                                                                                                              • DataOps is a term that has been co-opted and overloaded. What are the concepts that you are trying to convey with that term in the context of SQLMesh?
                                                                                                                                                                              • What are the rough edges in existing toolchains/workflows that you are trying to address with SQLMesh?
                                                                                                                                                                                • How do those rough edges impact the productivity and effectiveness of teams using those
                                                                                                                                                                                • Can you describe how SQLMesh is implemented?
                                                                                                                                                                                  • How have the design and goals evolved since you first started working on it?
                                                                                                                                                                                  • What are the lessons that you have learned from dbt which have informed the design and functionality of SQLMesh?
                                                                                                                                                                                  • For teams who have already invested in dbt, what is the migration path from or integration with dbt?
                                                                                                                                                                                  • You have some built-in integration with/awareness of orchestrators (currently Airflow). What are the benefits of making the transformation tool aware of the orchestrator?
                                                                                                                                                                                  • What do you see as the potential benefits of integration with e.g. data-diff?
                                                                                                                                                                                  • What are the second-order benefits of using a tool such as SQLMesh that addresses the more mechanical aspects of managing transformation workfows and the associated dependency chains?
                                                                                                                                                                                  • What are the most interesting, innovative, or unexpected ways that you have seen SQLMesh used?
                                                                                                                                                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on SQLMesh?
                                                                                                                                                                                  • When is SQLMesh the wrong choice?
                                                                                                                                                                                  • What do you have planned for the future of SQLMesh?
                                                                                                                                                                                  • Contact Info
                                                                                                                                                                                    • tobymao on GitHub
                                                                                                                                                                                    • @captaintobs on Twitter
                                                                                                                                                                                    • Website
                                                                                                                                                                                    • Parting Question
                                                                                                                                                                                      • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                                      • Closing Announcements
                                                                                                                                                                                        • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                                        • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                                        • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                                        • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                                                                                        • Links
                                                                                                                                                                                          • SQLMesh
                                                                                                                                                                                          • Tobiko Data
                                                                                                                                                                                          • SAS
                                                                                                                                                                                          • AirBnB Minerva
                                                                                                                                                                                          • SQLGlot
                                                                                                                                                                                          • Cron
                                                                                                                                                                                          • AST == Abstract Syntax Tree
                                                                                                                                                                                          • Pandas
                                                                                                                                                                                          • Terraform
                                                                                                                                                                                          • dbt
                                                                                                                                                                                            • Podcast Episode
                                                                                                                                                                                            • SQLFluff
                                                                                                                                                                                              • Podcast.__init__ Episode
                                                                                                                                                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                                                Sponsored By:

                                                                                                                                                                                                • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                                                                                RudderStack provides all your customer data pipelines in one platform. You can collect, transform, and route data across your entire stack with its event streaming, ETL, and reverse ETL pipelines.
                                                                                                                                                                                                RudderStack’s warehouse-first approach means it does not store sensitive information, and it allows you to leverage your existing data warehouse/data lake infrastructure to build a single source of truth for every team.
                                                                                                                                                                                                RudderStack also supports real-time use cases. You can Implement RudderStack SDKs once, then automatically send events to your warehouse and 150+ business tools, and you’ll never have to worry about API changes again.
                                                                                                                                                                                                Visit [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack) to sign up for free today, and snag a free T-Shirt just for being a Data Engineering Podcast listener.

                                                                                                                                                                                                Support Data Engineering Podcast

                                                                                                                                                                                                51 min
                                                                                                                                                                                              • How Column-Aware Development Tooling Yields Better Data Models
                                                                                                                                                                                                Summary

                                                                                                                                                                                                Architectural decisions are all based on certain constraints and a desire to optimize for different outcomes. In data systems one of the core architectural exercises is data modeling, which can have significant impacts on what is and is not possible for downstream use cases. By incorporating column-level lineage in the data modeling process it encourages a more robust and well-informed design. In this episode Satish Jayanthi explores the benefits of incorporating column-aware tooling in the data modeling process.

                                                                                                                                                                                                Announcements
                                                                                                                                                                                                • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                                                                • RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control. Their SDKs make event streaming from any app or website easy, and their extensive library of integrations enable you to automatically send data to hundreds of downstream tools. Sign up free at dataengineeringpodcast.com/rudderstack-
                                                                                                                                                                                                • Your host is Tobias Macey and today I'm interviewing Satish Jayanthi about the practice and promise of building a column-aware data architecture through intentional modeling
                                                                                                                                                                                                • Interview
                                                                                                                                                                                                  • Introduction
                                                                                                                                                                                                  • How did you get involved in the area of data management?
                                                                                                                                                                                                  • How has the move to the cloud for data warehousing/data platforms influenced the practice of data modeling?
                                                                                                                                                                                                    • There are ongoing conversations about the continued merits of dimensional modeling techniques in modern warehouses. What are the modeling practices that you have found to be most useful in large and complex data environments?
                                                                                                                                                                                                    • Can you describe what you mean by the term column-aware in the context of data modeling/data architecture?
                                                                                                                                                                                                      • What are the capabilities that need to be built into a tool for it to be effectively column-aware?
                                                                                                                                                                                                      • What are some of the ways that tools like dbt miss the mark in managing large/complex transformation workloads?
                                                                                                                                                                                                      • Column-awareness is obviously critical in the context of the warehouse. What are some of the ways that that information can be fed into other contexts? (e.g. ML, reverse ETL, etc.)
                                                                                                                                                                                                      • What is the importance of embedding column-level lineage awareness into transformation tool vs. layering on top w/ dedicated lineage/metadata tooling?
                                                                                                                                                                                                      • What are the most interesting, innovative, or unexpected ways that you have seen column-aware data modeling used?
                                                                                                                                                                                                      • What are the most interesting, unexpected, or challenging lessons that you have learned while working on building column-aware tooling?
                                                                                                                                                                                                      • When is column-aware modeling the wrong choice?
                                                                                                                                                                                                      • What are some additional resources that you recommend for individuals/teams who want to learn more about data modeling/column aware principles?
                                                                                                                                                                                                      • Contact Info
                                                                                                                                                                                                        • LinkedIn
                                                                                                                                                                                                        • Parting Question
                                                                                                                                                                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                                                          • Closing Announcements
                                                                                                                                                                                                            • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                                                            • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                                                            • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                                                            • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                                                                                                            • Links
                                                                                                                                                                                                              • Coalesce
                                                                                                                                                                                                                • Podcast Episode
                                                                                                                                                                                                                • Star Schema
                                                                                                                                                                                                                • Conformed Dimensions
                                                                                                                                                                                                                • Data Vault
                                                                                                                                                                                                                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                                                                  Sponsored By:

                                                                                                                                                                                                                  • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                                                                                                  RudderStack provides all your customer data pipelines in one platform. You can collect, transform, and route data across your entire stack with its event streaming, ETL, and reverse ETL pipelines.
                                                                                                                                                                                                                  RudderStack’s warehouse-first approach means it does not store sensitive information, and it allows you to leverage your existing data warehouse/data lake infrastructure to build a single source of truth for every team.
                                                                                                                                                                                                                  RudderStack also supports real-time use cases. You can Implement RudderStack SDKs once, then automatically send events to your warehouse and 150+ business tools, and you’ll never have to worry about API changes again.
                                                                                                                                                                                                                  Visit [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack) to sign up for free today, and snag a free T-Shirt just for being a Data Engineering Podcast listener.

                                                                                                                                                                                                                  Support Data Engineering Podcast

                                                                                                                                                                                                                  47 min
                                                                                                                                                                                                                • Build Better Tests For Your dbt Projects With Datafold And data-diff
                                                                                                                                                                                                                  Summary

                                                                                                                                                                                                                  Data engineering is all about building workflows, pipelines, systems, and interfaces to provide stable and reliable data. Your data can be stable and wrong, but then it isn't reliable. Confidence in your data is achieved through constant validation and testing. Datafold has invested a lot of time into integrating with the workflow of dbt projects to add early verification that the changes you are making are correct. In this episode Gleb Mezhanskiy shares some valuable advice and insights into how you can build reliable and well-tested data assets with dbt and data-diff.

                                                                                                                                                                                                                  Announcements
                                                                                                                                                                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                                                                                  • RudderStack helps you build a customer data platform on your warehouse or data lake. Instead of trapping data in a black box, they enable you to easily collect customer data from the entire stack and build an identity graph on your warehouse, giving you full visibility and control. Their SDKs make event streaming from any app or website easy, and their extensive library of integrations enable you to automatically send data to hundreds of downstream tools. Sign up free at dataengineeringpodcast.com/rudderstack
                                                                                                                                                                                                                  • Your host is Tobias Macey and today I'm interviewing Gleb Mezhanskiy about how to test your dbt projects with Datafold
                                                                                                                                                                                                                  • Interview
                                                                                                                                                                                                                    • Introduction
                                                                                                                                                                                                                    • How did you get involved in the area of data management?
                                                                                                                                                                                                                    • Can you describe what Datafold is and what's new since we last spoke? (July 2021 and July 2022 about data-diff)
                                                                                                                                                                                                                    • What are the roadblocks to data testing/validation that you see teams run into most often?
                                                                                                                                                                                                                      • How does the tooling used contribute to/help address those roadblocks?
                                                                                                                                                                                                                      • What are some of the error conditions/failure modes that data-diff can help identify in a dbt project?
                                                                                                                                                                                                                        • What are some examples of tests that need to be implemented by the engineer?
                                                                                                                                                                                                                        • In your experience working with data teams, what typically constitutes the "staging area" for a dbt project? (e.g. separate warehouse, namespaced tables, snowflake data copies, lakefs, etc.)
                                                                                                                                                                                                                        • Given a dbt project that is well tested and has data-diff as part of the validation suite, what are the challenges that teams face in managing the feedback cycle of running those tests?
                                                                                                                                                                                                                        • In application development there is the idea of the "testing pyramid", consisting of unit tests, integration tests, system tests, etc. What are the parallels to that in data projects?
                                                                                                                                                                                                                          • What are the limitations of the data ecosystem that make testing a bigger challenge than it might otherwise be?
                                                                                                                                                                                                                          • Beyond test execution, what are the other aspects of data health that need to be included in the development and deployment workflow of dbt projects? (e.g. freshness, time to delivery, etc.)
                                                                                                                                                                                                                          • What are the most interesting, innovative, or unexpected ways that you have seen Datafold and/or data-diff used for testing dbt projects?
                                                                                                                                                                                                                          • What are the most interesting, unexpected, or challenging lessons that you have learned while working on dbt testing internally or with your customers?
                                                                                                                                                                                                                          • When is Datafold/data-diff the wrong choice for dbt projects?
                                                                                                                                                                                                                          • What do you have planned for the future of Datafold?
                                                                                                                                                                                                                          • Contact Info
                                                                                                                                                                                                                            • LinkedIn
                                                                                                                                                                                                                            • Closing Announcements
                                                                                                                                                                                                                              • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                                                                              • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                                                                              • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                                                                              • To help other people find the show please leave a review on Apple Podcasts and tell your friends and co-workers
                                                                                                                                                                                                                              • Parting Question
                                                                                                                                                                                                                                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                                                                                • Links
                                                                                                                                                                                                                                  • Datafold
                                                                                                                                                                                                                                    • Podcast Episode
                                                                                                                                                                                                                                    • data-diff
                                                                                                                                                                                                                                      • Podcast Episode
                                                                                                                                                                                                                                      • dbt
                                                                                                                                                                                                                                      • Dagster
                                                                                                                                                                                                                                      • dbt-cloud slim CI
                                                                                                                                                                                                                                      • GitHub Actions
                                                                                                                                                                                                                                      • Jenkins
                                                                                                                                                                                                                                      • Circle CI
                                                                                                                                                                                                                                      • Dolt
                                                                                                                                                                                                                                      • Malloy
                                                                                                                                                                                                                                      • LakeFS
                                                                                                                                                                                                                                      • Planetscale
                                                                                                                                                                                                                                      • Snowflake Zero Copy Cloning
                                                                                                                                                                                                                                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                                                                                        Special Guest: Gleb Mezhanskiy.

                                                                                                                                                                                                                                        Sponsored By:

                                                                                                                                                                                                                                        • Rudderstack: ![Rudderstack](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/CKNV8HZ6.png)
                                                                                                                                                                                                                                        RudderStack provides all your customer data pipelines in one platform. You can collect, transform, and route data across your entire stack with its event streaming, ETL, and reverse ETL pipelines.
                                                                                                                                                                                                                                        RudderStack’s warehouse-first approach means it does not store sensitive information, and it allows you to leverage your existing data warehouse/data lake infrastructure to build a single source of truth for every team.
                                                                                                                                                                                                                                        RudderStack also supports real-time use cases. You can Implement RudderStack SDKs once, then automatically send events to your warehouse and 150+ business tools, and you’ll never have to worry about API changes again.
                                                                                                                                                                                                                                        Visit [dataengineeringpodcast.com/rudderstack](https://www.dataengineeringpodcast.com/rudderstack) to sign up for free today, and snag a free T-Shirt just for being a Data Engineering Podcast listener.

                                                                                                                                                                                                                                        Support Data Engineering Podcast

                                                                                                                                                                                                                                        49 min

                                                                                                                                                                                                                                      About Data Engineering Podcast

                                                                                                                                                                                                                                      From the publisher's feed

                                                                                                                                                                                                                                      This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

                                                                                                                                                                                                                                      More shows like Data Engineering Podcast

                                                                                                                                                                                                                                      This Week in Startups by Jason Calacanis

                                                                                                                                                                                                                                      This Week in Startups

                                                                                                                                                                                                                                      1,290 Listeners

                                                                                                                                                                                                                                      The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                                                                                                                      The Changelog: Software Development, Open Source

                                                                                                                                                                                                                                      286 Listeners

                                                                                                                                                                                                                                      The a16z Show by Andreessen Horowitz

                                                                                                                                                                                                                                      The a16z Show

                                                                                                                                                                                                                                      1,087 Listeners

                                                                                                                                                                                                                                      Software Engineering Daily by Software Engineering Daily

                                                                                                                                                                                                                                      Software Engineering Daily

                                                                                                                                                                                                                                      623 Listeners

                                                                                                                                                                                                                                      Risky Business by Risky Business Media

                                                                                                                                                                                                                                      Risky Business

                                                                                                                                                                                                                                      375 Listeners

                                                                                                                                                                                                                                      Talk Python To Me by Michael Kennedy

                                                                                                                                                                                                                                      Talk Python To Me

                                                                                                                                                                                                                                      582 Listeners

                                                                                                                                                                                                                                      Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                                                                                                                      Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                                                                                                                      305 Listeners

                                                                                                                                                                                                                                      NVIDIA AI Podcast by NVIDIA

                                                                                                                                                                                                                                      NVIDIA AI Podcast

                                                                                                                                                                                                                                      338 Listeners

                                                                                                                                                                                                                                      Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                                                                                                                      Syntax - Tasty Web Development Treats

                                                                                                                                                                                                                                      985 Listeners

                                                                                                                                                                                                                                      Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                                                                                                                      Practical AI

                                                                                                                                                                                                                                      203 Listeners

                                                                                                                                                                                                                                      Dwarkesh Podcast by Dwarkesh Patel

                                                                                                                                                                                                                                      Dwarkesh Podcast

                                                                                                                                                                                                                                      565 Listeners

                                                                                                                                                                                                                                      The Data Engineering Show by The Firebolt Data Bros

                                                                                                                                                                                                                                      The Data Engineering Show

                                                                                                                                                                                                                                      8 Listeners

                                                                                                                                                                                                                                      Latent Space: The AI Engineer Podcast by Latent.Space

                                                                                                                                                                                                                                      Latent Space: The AI Engineer Podcast

                                                                                                                                                                                                                                      102 Listeners

                                                                                                                                                                                                                                      This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                                                                                                                                                                                                                                      This Day in AI Podcast

                                                                                                                                                                                                                                      222 Listeners

                                                                                                                                                                                                                                      The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                                                                                                                                                                                                                                      The AI Daily Brief: Artificial Intelligence News and Analysis

                                                                                                                                                                                                                                      684 Listeners