Data Engineering Podcast

Defining DataOps with Chris Bergh - Episode 26


Listen Later

Summary

Managing an analytics project can be difficult due to the number of systems involved and the need to ensure that new information can be delivered quickly and reliably. That challenge can be met by adopting practices and principles from lean manufacturing and agile software development, and the cross-functional collaboration, feedback loops, and focus on automation in the DevOps movement. In this episode Christopher Bergh discusses ways that you can start adding reliability and speed to your workflow to deliver results with confidence and consistency.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
  • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
  • Your host is Tobias Macey and today I’m interviewing Christopher Bergh about DataKitchen and the rise of DataOps
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • How do you define DataOps?
      • How does it compare to the practices encouraged by the DevOps movement?
      • How does it relate to or influence the role of a data engineer?

      • How does a DataOps oriented workflow differ from other existing approaches for building data platforms?

      • One of the aspects of DataOps that you call out is the practice of providing multiple environments to provide a platform for testing the various aspects of the analytics workflow in a non-production context. What are some of the techniques that are available for managing data in appropriate volumes across those deployments?

      • The practice of testing logic as code is fairly well understood and has a large set of existing tools. What have you found to be some of the most effective methods for testing data as it flows through a system?

      • One of the practices of DevOps is to create feedback loops that can be used to ensure that business needs are being met. What are the metrics that you track in your platform to define the value that is being created and how the various steps in the workflow are proceeding toward that goal?

        • In order to keep feedback loops fast it is necessary for tests to run quickly. How do you balance the need for larger quantities of data to be used for verifying scalability/performance against optimizing for cost and speed in non-production environments?

        • How does the DataKitchen platform simplify the process of operationalizing a data analytics workflow?

        • As the need for rapid iteration and deployment of systems to capture, store, process, and analyze data becomes more prevalent how do you foresee that feeding back into the ways that the landscape of data tools are designed and developed?

        • Contact Info
          • LinkedIn
          • @ChrisBergh on Twitter
          • Email
          • Parting Question
            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
            • Links
              • DataOps Manifesto
              • DataKitchen
              • 2017: The Year Of DataOps
              • Air Traffic Control
              • Chief Data Officer (CDO)
              • Gartner
              • W. Edwards Deming
              • DevOps
              • Total Quality Management (TQM)
              • Informatica
              • Talend
              • Agile Development
              • Cattle Not Pets
              • IDE (Integrated Development Environment)
              • Tableau
              • Delphix
              • Dremio
              • Pachyderm
              • Continuous Delivery by Jez Humble and Dave Farley
              • SLAs (Service Level Agreements)
              • XKCD Image Recognition Comic
              • Airflow
              • Luigi
              • DataKitchen Documentation
              • Continuous Integration
              • Continous Delivery
              • Docker
              • Version Control
              • Git
              • Looker
              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                Support Data Engineering Podcast

                ...more
                View all episodesView all episodes
                Download on the App Store

                Data Engineering PodcastBy Tobias Macey

                • 4.6
                • 4.6
                • 4.6
                • 4.6
                • 4.6

                4.6

                135 ratings


                More shows like Data Engineering Podcast

                View all
                Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                Software Engineering Radio - the podcast for professional software developers

                272 Listeners

                The Changelog: Software Development, Open Source by Changelog Media

                The Changelog: Software Development, Open Source

                283 Listeners

                The Cloudcast by Massive Studios

                The Cloudcast

                153 Listeners

                Thoughtworks Technology Podcast by Thoughtworks

                Thoughtworks Technology Podcast

                41 Listeners

                Data Skeptic by Kyle Polich

                Data Skeptic

                483 Listeners

                Talk Python To Me by Michael Kennedy

                Talk Python To Me

                592 Listeners

                Software Engineering Daily by Software Engineering Daily

                Software Engineering Daily

                624 Listeners

                The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

                The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

                444 Listeners

                Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                Super Data Science: ML & AI Podcast with Jon Krohn

                298 Listeners

                Python Bytes by Michael Kennedy and Brian Okken

                Python Bytes

                213 Listeners

                DataFramed by DataCamp

                DataFramed

                266 Listeners

                Practical AI by Practical AI LLC

                Practical AI

                190 Listeners

                The Stack Overflow Podcast by The Stack Overflow Podcast

                The Stack Overflow Podcast

                64 Listeners

                The Real Python Podcast by Real Python

                The Real Python Podcast

                140 Listeners

                Latent Space: The AI Engineer Podcast by swyx + Alessio

                Latent Space: The AI Engineer Podcast

                77 Listeners