Data Engineering Podcast

Brief Conversations From The Open Data Science Conference: Part 1 - Episode 30


Listen Later

Summary

The Open Data Science Conference brings together a variety of data professionals each year in Boston. This week’s episode consists of a pair of brief interviews conducted on-site at the conference. First up you’ll hear from Alan Anders, the CTO of Applecart about their challenges with getting Spark to scale for constructing an entity graph from multiple data sources. Next I spoke with Stepan Pushkarev, the CEO, CTO, and Co-Founder of Hydrosphere.io about the challenges of running machine learning models in production and how his team tracks key metrics and samples production data to re-train and re-deploy those models for better accuracy and more robust operation.

Preamble
  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
  • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
  • Your host is Tobias Macey and this week I attended the Open Data Science Conference in Boston and recorded a few brief interviews on-site. First up you’ll hear from Alan Anders, the CTO of Applecart about their challenges with getting Spark to scale for constructing an entity graph from multiple data sources. Next I spoke with Stepan Pushkarev, the CEO, CTO, and Co-Founder of Hydrosphere.io about the challenges of running machine learning models in production and how his team tracks key metrics and samples production data to re-train and re-deploy those models for better accuracy and more robust operation.
  • Interview
    Alan Anders from Applecart
    • What are the challenges of gathering and processing data from multiple data sources and representing them in a unified manner for merging into single entities?
    • What are the biggest technical hurdles at Applecart?
    • Contact Info
      • @alanjanders on Twitter
      • LinkedIn
      • Parting Question
        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
        • Links
          • Spark
          • DataBricks
          • DataBricks Delta
          • Applecart
          • Stepan Pushkarev from Hydrosphere.io
            • What is Hydropshere.io?
            • What metrics do you track to determine when a machine learning model is not producing an appropriate output?
            • How do you determine which data points to sample for retraining the model?
            • How does the role of a machine learning engineer differ from data engineers and data scientists?
            • Contact Info
              • LinkedIn
              • Parting Question
                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                • Links
                  • Hydrosphere
                  • Machine Learning Engineer
                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                    Support Data Engineering Podcast

                    ...more
                    View all episodesView all episodes
                    Download on the App Store

                    Data Engineering PodcastBy Tobias Macey

                    • 4.5
                    • 4.5
                    • 4.5
                    • 4.5
                    • 4.5

                    4.5

                    142 ratings


                    More shows like Data Engineering Podcast

                    View all
                    The Changelog: Software Development, Open Source by Changelog Media

                    The Changelog: Software Development, Open Source

                    289 Listeners

                    Software Engineering Daily by Software Engineering Daily

                    Software Engineering Daily

                    624 Listeners

                    Talk Python To Me by Michael Kennedy

                    Talk Python To Me

                    583 Listeners

                    Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                    Super Data Science: ML & AI Podcast with Jon Krohn

                    302 Listeners

                    NVIDIA AI Podcast by NVIDIA

                    NVIDIA AI Podcast

                    343 Listeners

                    Practical AI by Practical AI LLC

                    Practical AI

                    204 Listeners

                    AWS Podcast by Amazon Web Services

                    AWS Podcast

                    205 Listeners

                    Last Week in AI by Skynet Today

                    Last Week in AI

                    305 Listeners

                    Dwarkesh Podcast by Dwarkesh Patel

                    Dwarkesh Podcast

                    523 Listeners

                    The Data Engineering Show by The Firebolt Data Bros

                    The Data Engineering Show

                    8 Listeners

                    No Priors: Artificial Intelligence | Technology | Startups by Conviction

                    No Priors: Artificial Intelligence | Technology | Startups

                    129 Listeners

                    Latent Space: The AI Engineer Podcast by swyx + Alessio

                    Latent Space: The AI Engineer Podcast

                    92 Listeners

                    This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                    This Day in AI Podcast

                    227 Listeners

                    The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                    The AI Daily Brief: Artificial Intelligence News and Analysis

                    633 Listeners

                    AI + a16z by a16z

                    AI + a16z

                    36 Listeners