Data Engineering Podcast

Data Orchestration For Hybrid Cloud Analytics


Listen Later

Summary

The scale and complexity of the systems that we build to satisfy business requirements is increasing as the available tools become more sophisticated. In order to bridge the gap between legacy infrastructure and evolving use cases it is necessary to create a unifying set of components. In this episode Dipti Borkar explains how the emerging category of data orchestration tools fills this need, some of the existing projects that fit in this space, and some of the ways that they can work together to simplify projects such as cloud migration and hybrid cloud environments. It is always useful to get a broad view of new trends in the industry and this was a helpful perspective on the need to provide mechanisms to decouple physical storage from computing capacity.

Announcements
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline, or want to test out the projects you hear about on the show, you’ll need somewhere to deploy it, so check out our friends at Linode. With 200Gbit private networking, scalable shared block storage, and a 40Gbit public network, you’ve got everything you need to run a fast, reliable, and bullet-proof data platform. If you need global distribution, they’ve got that covered too with world-wide datacenters including new ones in Toronto and Mumbai. And for your machine learning workloads, they just announced dedicated CPU instances. Go to dataengineeringpodcast.com/linode today to get a $20 credit and launch a new server in under a minute. And don’t forget to thank them for their continued support of this show!
  • This week’s episode is also sponsored by Datacoral, an AWS-native, serverless, data infrastructure that installs in your VPC. Datacoral helps data engineers build and manage the flow of data pipelines without having to manage any infrastructure, meaning you can spend your time invested in data transformations and business needs, rather than pipeline maintenance. Raghu Murthy, founder and CEO of Datacoral built data infrastructures at Yahoo! and Facebook, scaling from terabytes to petabytes of analytic data. He started Datacoral with the goal to make SQL the universal data programming language. Visit dataengineeringpodcast.com/datacoral today to find out more.
  • You listen to this show to learn and stay up to date with what’s happening in databases, streaming platforms, big data, and everything else you need to know about modern data management. For even more opportunities to meet, listen, and learn from your peers you don’t want to miss out on this year’s conference season. We have partnered with organizations such as O’Reilly Media, Dataversity, Corinium Global Intelligence, Alluxio, and Data Council. Upcoming events include the combined events of the Data Architecture Summit and Graphorum, the Data Orchestration Summit, and Data Council in NYC. Go to dataengineeringpodcast.com/conferences to learn more about these and other events, and take advantage of our partner discounts to save money when you register today.
  • Your host is Tobias Macey and today I’m interviewing Dipti Borkark about data orchestration and how it helps in migrating data workloads to the cloud
  • Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by describing what you mean by the term "Data Orchestration"?
      • How does it compare to the concept of "Data Virtualization"?
      • What are some of the tools and platforms that fit under that umbrella?
      • What are some of the motivations for organizations to use the cloud for their data oriented workloads?
        • What are they giving up by using cloud resources in place of on-premises compute?
        • For businesses that have invested heavily in their own datacenters, what are some ways that they can begin to replicate some of the benefits of cloud environments?
        • What are some of the common patterns for cloud migration projects and what challenges do they present?
          • Do you have advice on useful metrics to track for determining project completion or success criteria?
          • How do businesses approach employee education for designing and implementing effective systems for achieving their migration goals?
          • Can you talk through some of the ways that different data orchestration tools can be composed together for a cloud migration effort?
            • What are some of the common pain points that organizations encounter when working on hybrid implementations?
            • What are some of the missing pieces in the data orchestration landscape?
              • Are there any efforts that you are aware of that are aiming to fill those gaps?
              • Where is the data orchestration market heading, and what are some industry trends that are driving it?
                • What projects are you most interested in or excited by?
                • For someone who wants to learn more about data orchestration and the benefits the technologies can provide, what are some resources that you would recommend?
                • Contact Info
                  • LinkedIn
                  • @dborkar on Twitter
                  • Parting Question
                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                    • Closing Announcements
                      • Thank you for listening! Don’t forget to check out our other show, Podcast.__init__ to learn about the Python language, its community, and the innovative ways it is being used.
                      • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                      • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                      • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                      • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                      • Links
                        • Alluxio
                          • Podcast Episode
                          • UC San Diego
                          • Couchbase
                          • Presto
                            • Podcast Episode
                            • Spark SQL
                            • Data Orchestration
                            • Data Virtualization
                            • PyTorch
                              • Podcast.init Episode
                              • Rook storage orchestration
                              • PySpark
                              • MinIO
                                • Podcast Episode
                                • Kubernetes
                                • Openstack
                                • Hadoop
                                • HDFS
                                • Parquet Files
                                  • Podcast Episode
                                  • ORC Files
                                  • Hive Metastore
                                  • Iceberg Table Format
                                    • Podcast Episode
                                    • Data Orchestration Summit
                                    • Star Schema
                                    • Snowflake Schema
                                    • Data Warehouse
                                    • Data Lake
                                    • Teradata
                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                      Support Data Engineering Podcast

                                      ...more
                                      View all episodesView all episodes
                                      Download on the App Store

                                      Data Engineering PodcastBy Tobias Macey

                                      • 4.6
                                      • 4.6
                                      • 4.6
                                      • 4.6
                                      • 4.6

                                      4.6

                                      135 ratings


                                      More shows like Data Engineering Podcast

                                      View all
                                      Software Engineering Radio - the podcast for professional software developers by se-radio@computer.org

                                      Software Engineering Radio - the podcast for professional software developers

                                      272 Listeners

                                      The Changelog: Software Development, Open Source by Changelog Media

                                      The Changelog: Software Development, Open Source

                                      283 Listeners

                                      The Cloudcast by Massive Studios

                                      The Cloudcast

                                      152 Listeners

                                      Thoughtworks Technology Podcast by Thoughtworks

                                      Thoughtworks Technology Podcast

                                      41 Listeners

                                      Data Skeptic by Kyle Polich

                                      Data Skeptic

                                      482 Listeners

                                      Talk Python To Me by Michael Kennedy

                                      Talk Python To Me

                                      592 Listeners

                                      Software Engineering Daily by Software Engineering Daily

                                      Software Engineering Daily

                                      625 Listeners

                                      The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence) by Sam Charrington

                                      The TWIML AI Podcast (formerly This Week in Machine Learning & Artificial Intelligence)

                                      443 Listeners

                                      Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                      Super Data Science: ML & AI Podcast with Jon Krohn

                                      296 Listeners

                                      Python Bytes by Michael Kennedy and Brian Okken

                                      Python Bytes

                                      213 Listeners

                                      DataFramed by DataCamp

                                      DataFramed

                                      266 Listeners

                                      Practical AI by Practical AI LLC

                                      Practical AI

                                      189 Listeners

                                      The Stack Overflow Podcast by The Stack Overflow Podcast

                                      The Stack Overflow Podcast

                                      64 Listeners

                                      The Real Python Podcast by Real Python

                                      The Real Python Podcast

                                      140 Listeners

                                      Latent Space: The AI Engineer Podcast by swyx + Alessio

                                      Latent Space: The AI Engineer Podcast

                                      77 Listeners