Data Engineering Podcast

Advice On Scaling Your Data Pipeline Alongside Your Business with Christian Heinzmann - Episode 61


Listen Later

Summary

Every business needs a pipeline for their critical data, even if it is just pasting into a spreadsheet. As the organization grows and gains more customers, the requirements for that pipeline will change. In this episode Christian Heinzmann, Head of Data Warehousing at Grubhub, discusses the various requirements for data pipelines and how the overall system architecture evolves as more data is being processed. He also covers the changes in how the output of the pipelines are used, how that impacts the expectations for accuracy and availability, and some useful advice on build vs. buy for the components of a data platform.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
  • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
  • Your host is Tobias Macey and today I’m interviewing Christian Heinzmann about how data pipelines evolve as your business grows
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by sharing your definition of a data pipeline?
      • At what point in the life of a project or organization should you start thinking about building a pipeline?

      • In the early stages when the scale of the data and business are still small, what are some of the design characteristics that you should be targeting for your pipeline?

        • What metrics/use cases should you be optimizing for at this point?

        • What are some of the indicators that you look for to signal that you are reaching the next order of magnitude in terms of scale?

          • How do the design requirements for a data pipeline change as you reach this stage?
          • What are some of the challenges and complexities that begin to present themselves as you build and run your pipeline at medium scale?

          • What are some of the changes that are necessary as you move to a large scale data pipeline?

          • At each level of scale it is important to minimize the impact of the ETL process on the source systems. What are some strategies that you have employed to avoid degrading the performance of the application systems?

          • In recent years there has been a shift to using data lakes as a staging ground before performing transformations. What are your thoughts on that approach?

          • When performing transformations there is a potential for discarding information or losing fidelity. How have you worked to reduce the impact of this effect?

          • Transformations of the source data can be brittle when the format or volume changes. How do you design the pipeline to be resilient to these types of changes?

          • What are your selection criteria when determining what workflow or ETL engines to use in your pipeline?

            • How has your preference of build vs buy changed at different scales of operation and as new/different projects become available?

            • What are some of the dead ends or edge cases that you have had to deal with in your current role at Grubhub?

            • What are some of the common mistakes or overlooked aspects of building a data pipeline that you have seen?

            • What are your plans for improving your current pipeline at Grubhub?

            • What are some references that you recommend for anyone who is designing a new data platform?

            • Contact Info
              • @sirchristian on Twitter
              • Blog
              • sirchristian on GitHub
              • Parting Question
                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                • Links
                  • Scaling ETL blog post
                  • GrubHub
                  • Data Warehouse
                  • Redshift
                  • Spark
                    • Spark In Action Podcast Episode

                    • Hive

                    • Amazon EMR

                    • Looker

                      • Podcast Episode

                      • Redash

                      • Metabase

                        • Podcast Episode

                        • A Primer on Enterprise Data Curation

                        • Pub/Sub (Publish-Subscribe Pattern)

                        • Change Data Capture

                        • Jenkins

                        • Python

                        • Azkaban

                        • Luigi

                        • Zendesk

                        • Data Lineage

                        • AirBnB Engineering Blog

                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                          Support Data Engineering Podcast

                          ...more
                          View all episodesView all episodes
                          Download on the App Store

                          Data Engineering PodcastBy Tobias Macey

                          • 4.6
                          • 4.6
                          • 4.6
                          • 4.6
                          • 4.6

                          4.6

                          135 ratings


                          More shows like Data Engineering Podcast

                          View all
                          Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                          Software Engineering Radio - the podcast for professional software developers

                          272 Listeners

                          The Changelog: Software Development, Open Source by Changelog Media

                          The Changelog: Software Development, Open Source

                          283 Listeners

                          The Cloudcast by Massive Studios

                          The Cloudcast

                          152 Listeners

                          Thoughtworks Technology Podcast by Thoughtworks

                          Thoughtworks Technology Podcast

                          41 Listeners

                          Data Skeptic by Kyle Polich

                          Data Skeptic

                          482 Listeners

                          Talk Python To Me by Michael Kennedy

                          Talk Python To Me

                          592 Listeners

                          Software Engineering Daily by Software Engineering Daily

                          Software Engineering Daily

                          624 Listeners

                          The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

                          The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

                          443 Listeners

                          Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                          Super Data Science: ML & AI Podcast with Jon Krohn

                          298 Listeners

                          Python Bytes by Michael Kennedy and Brian Okken

                          Python Bytes

                          213 Listeners

                          DataFramed by DataCamp

                          DataFramed

                          266 Listeners

                          Practical AI by Practical AI LLC

                          Practical AI

                          189 Listeners

                          The Stack Overflow Podcast by The Stack Overflow Podcast

                          The Stack Overflow Podcast

                          64 Listeners

                          The Real Python Podcast by Real Python

                          The Real Python Podcast

                          140 Listeners

                          Latent Space: The AI Engineer Podcast by swyx + Alessio

                          Latent Space: The AI Engineer Podcast

                          77 Listeners