The Python Podcast.__init__

The Python Podcast.__init__

By Tobias MaceyTechnologyEducation
Download on the App Store

The Python Podcast.__init__ episodes

  • Digging Into Dagster: An Opinionated Open Source Framework For Data Orchestration
    Summary

    Data applications are complex and continually evolving, often requiring collaboration across multiple teams. In order to keep everyone on the same page a high level abstraction is needed to facilitate a cross-cutting view of the data orchestration across integration, transformation, analytics, and machine learning. Dagster is an innovative new framework that leans on the power and flexibility of Python to provide an extensible interface to the complete lifecycle of data projects. In this episode Nick Schrock explains how he designed the Dagster project to allow for integration with the entire data ecosystem while providing an opinionated structure for connecting the different stages of computation. He also discusses how he is working to grow an open ecosystem around the Dagster project, and his thoughts on building a sustainable business on top of it without compromising the integrity of the community. This was a great conversation about playing the long game when building a business while providing a valuable utility to a complex problem domain.

    Announcements
    • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
    • This portion of Python Podcast is brought to you by Datadog. Do you have an app in production that is slower than you like? Is its performance all over the place (sometimes fast, sometimes slow)? Do you know why? With Datadog, you will. You can troubleshoot your app’s performance with Datadog’s end-to-end tracing and in one click correlate those Python traces with related logs and metrics. Use their detailed flame graphs to identify bottlenecks and latency in that app of yours. Start tracking the performance of your apps with a free trial at pythonpodcast.com/datadog. If you sign up for a trial and install the agent, Datadog will send you a free t-shirt.
    • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
    • Your host as usual is Tobias Macey and today I’m interviewing Nick Schrock about Dagster, an open source data orchestrator for powering data engineering, analytics, and machine learning
    • Interview
      • Introductions
      • How did you get introduced to Python?
      • Can you start by describing what Dagster is and how it got started?
      • What are the most common difficulties that organizations face when working with data projects?
        • How does Dagster help in addressing those challenges?
        • There are a number of workflow orchestration platforms, spanning a few generations of tooling. What do you see as the defining characteristics of the various options, and how does Dagster fit in that ecosystem?
        • What are the assumptions that you made at the start of building Dagster and how have they been challenged, updated, or invalidated over the past year of working with end users?
        • How are the internals of Dagster implemented?
          • How has the design changed or evolved since you first began working on it?
          • For someone who is building on top of Dagster, what is their workflow from first steps through to production?
          • What are your guiding principles for desigining the user facing API?
          • What are the available extension points for Dagster?
          • What was your reason for implementing Dagster as a Python framework?
            • With the benefit of hindsight, would you make the same decision today?
            • What are some of the most interesting, innovative, or unexpected ways that you have seen Dagster used?
            • What are the most interesting, unexpected, or challenging lessons that you have learned while building Dagster and working to grow its ecosystem?
            • When is Dagster the wrong choice?
            • As you continue to build Dagster, what is your vision for it and its ecosystem?
              • What are the next steps that you are taking to achieve that vision?
              • Keep In Touch
                • @schrockn on Twitter
                • schrockn on GitHub
                • LinkedIn
                • Picks
                  • Tobias
                    • Caddy web server
                    • Nick
                      • Black code formatter
                      • Closing Announcements
                        • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
                        • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                        • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                        • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                        • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat
                        • Links
                          • Dagster
                          • Elementl
                          • IronPython
                          • Fluent Python
                          • GraphQL
                          • Maslow’s Hierarchy of Needs
                          • Hierarchy of Data Needs
                          • DAG == Directed Acyclic Graph
                          • Informatica
                          • Airflow
                          • Luigi
                          • Dagster Config Schema
                          • Dask
                            • Data Engineering Podcast Episode
                            • Coiled Episode
                            • gRPC
                            • MyPy
                              • Podcast Episode
                              • Data Lineage
                              • Pandas
                                • Podcast Episode
                                • Amundsen
                                  • Podcast Episode
                                  • DataHub
                                    • Podcast Episode
                                    • Gatsby.js
                                    • Panama Papers
                                    • Mode Analytics
                                      • Podcast Episode
                                      • Papermill
                                        • Podcast Episode
                                        • DBT
                                          • Podcast Episode
                                          • Databricks
                                          • Tobias’ Dagster Repository
                                          • The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

                                            1 hr
                                          • When, Why, and How To Use Web Scraping In A Nutshell
                                            The internet is a rich source of information, but a majority of it isn't accessible programmatically through APIs or databases. To address that shortcoming there are a variety of web scraping frameworks that aid in extracting structured data from web pages. In this episode Attila Tóth shares the challenges of web data extraction, the ways that you can use it, and how Scrapy and ScrapingHub can help you with your projects.
                                            42 min
                                          • When, Why, and How To Use Web Scraping In A Nutshell
                                            Summary

                                            The internet is a rich source of information, but a majority of it isn’t accessible programmatically through APIs or databases. To address that shortcoming there are a variety of web scraping frameworks that aid in extracting structured data from web pages. In this episode Attila Tóth shares the challenges of web data extraction, the ways that you can use it, and how Scrapy and ScrapingHub can help you with your projects.

                                            Announcements
                                            • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
                                            • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
                                            • This portion of Python Podcast is brought to you by Datadog. Do you have an app in production that is slower than you like? Is its performance all over the place (sometimes fast, sometimes slow)? Do you know why? With Datadog, you will. You can troubleshoot your app’s performance with Datadog’s end-to-end tracing and in one click correlate those Python traces with related logs and metrics. Use their detailed flame graphs to identify bottlenecks and latency in that app of yours. Start tracking the performance of your apps with a free trial at datadog.com/pythonpodcast. If you sign up for a trial and install the agent, Datadog will send you a free t-shirt.
                                            • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
                                            • Your host as usual is Tobias Macey and today I’m interviewing Attila Tóth about doing data extraction with web scraping.
                                            • Interview
                                              • Introductions
                                              • How did you get introduced to Python?
                                              • Can you start by explaining what web scraping is and when you might want to use it?
                                                • How did you first get started with web scraping?
                                                • There are a number of options for web scraping tools in Python, as well as other languages. What are the characteristics of the Scrapy project and community that have made it stand out and retain such widespread popularity?
                                                • One of the perpetual questions with web scraping is that of copyright and content ownership. What should we all be aware of when scraping a given website?
                                                • What are some of the most challenging aspects of crawling and scraping the web?
                                                  • What are some of the features of Scrapy that aid in those challenges?
                                                  • Once you have retrieved the content from a site, what are some of the considerations for storing and processing the data that we should be thinking about?
                                                  • How can we guard against a scraper breaking due to changes in the layout of a site, or simple updates that weren’t accounted for in the initial implementation?
                                                  • What are some of the most complicated aspects of scaling web scrapers?
                                                  • For someone who is interested in using Scrapy, what are some of the common pitfalls that they should be aware of?
                                                  • What are some of the most interesting, innovative, or unexpected projects that are built with Scrapy and ScrapingHub?
                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working with web scrapers and ScrapingHub?
                                                  • What resources would you recommend to anyone who is looking to learn more about web scraping?
                                                  • Keep In Touch
                                                    • LinkedIn
                                                    • Picks
                                                      • Tobias
                                                        • Gov’t Mule
                                                        • Attila
                                                          • Awesome Web Scraping
                                                          • Awesome Scrapy
                                                          • Closing Announcements
                                                            • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
                                                            • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                            • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                            • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                                                            • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat
                                                            • Links
                                                              • Web Scraping
                                                              • ScrapingHub
                                                              • Java
                                                              • Android
                                                              • Scrapy
                                                              • JSoup
                                                              • HTMLUnit
                                                              • Selenium
                                                              • Pandas
                                                              • robots.txt
                                                              • Puppeteer
                                                              • Splash
                                                              • The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

                                                                42 min
                                                              • Working In The Code Mines: Mining Software Repositories With PyDriller
                                                                A large portion of the software industry has standardized on Git as the version control sytem of choice. But have you thought about all of the information that you are generating with your branches, commits, and code changes? Davide Spadini created the PyDriller framework to simplify the work of mining software repositories to perform research on the technical and social aspects of software engineering. In this episode he shares some of the insights that you can gain by exploring the history of your code, the complexities of building a framework to interact with Git, and some of the interesting ways that PyDriller can be used to inform your own development practices.
                                                                41 min
                                                              • Working In The Code Mines: Mining Software Repositories With PyDriller
                                                                Summary

                                                                A large portion of the software industry has standardized on Git as the version control sytem of choice. But have you thought about all of the information that you are generating with your branches, commits, and code changes? Davide Spadini created the PyDriller framework to simplify the work of mining software repositories to perform research on the technical and social aspects of software engineering. In this episode he shares some of the insights that you can gain by exploring the history of your code, the complexities of building a framework to interact with Git, and some of the interesting ways that PyDriller can be used to inform your own development practices.

                                                                Announcements
                                                                • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
                                                                • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
                                                                • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
                                                                • Your host as usual is Tobias Macey and today I’m interviewing Davide Spadini about PyDriller, a framework for mining software repositories
                                                                • Interview
                                                                  • Introductions
                                                                  • How did you get introduced to Python?
                                                                  • Can you start by describing what PyDriller is and how the project got started?
                                                                    • How is Pydriller different from other Git frameworks?
                                                                    • What kinds of information can you discover by mining a software repository?
                                                                      • Where and how might the collected information be used?
                                                                      • What are the limitations of the capabilities offered by Git for investigating the repository?
                                                                      • What are the additional metrics that you are able to extract using PyDriller?
                                                                      • Can you describe how PyDriller itself is implemented?
                                                                        • How has the project evolved since you first began working on it?
                                                                        • I noticed that for testing PyDriller you crafted a set of repositories to serve as test cases. What has been the most complex or challenging aspect of writing meaningful tests to ensure a reasonable coverage of this problem domain?
                                                                        • What would be required to add support for other version control systems?
                                                                        • How have you used PyDriller in your own research?
                                                                        • What are some of the most interesting, unexpected, or innovative ways that you have seen PyDriller used?
                                                                        • What are some of the most interesting, unexpected, or challenging lessons that you have learned while working on and with PyDriller?
                                                                        • What do you have planned for the future of PyDriller?
                                                                        • Keep In Touch
                                                                          • Website
                                                                          • ishepard on GitHub
                                                                          • @DavideSpadini on Twitter
                                                                          • Picks
                                                                            • Tobias
                                                                              • pre-commit
                                                                              • Davide
                                                                                • Fall guys
                                                                                • Closing Announcements
                                                                                  • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
                                                                                  • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                  • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                  • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                                                                                  • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat
                                                                                  • Links
                                                                                    • PyDriller
                                                                                    • Delft
                                                                                    • Git
                                                                                    • GitPython
                                                                                    • PyGit2
                                                                                    • RepoDriller
                                                                                    • Mining Software Repositories Conference
                                                                                    • Lizard
                                                                                    • Hadoop
                                                                                    • Mercurial
                                                                                      • Podcast Episode
                                                                                      • Subversion
                                                                                      • CVS
                                                                                      • Neo4J
                                                                                      • GraphRepo
                                                                                      • The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

                                                                                        41 min
                                                                                      • Building The Open Data Ecosystem For Music And More At Metabrainz
                                                                                        The Musicbrainz project was an early entry in the movement to build an open data ecosystem. In recent years, the Metabrainz Foundation has fostered a growing ecosystem of projects to support the contribution of, and access to, metadata, listening habits, and review of music. The majority of those projects are written in Python, and in this episode Param Singh explains how they are built, how they fit together, and how they support the goals of the Metabrains Foundation. This was an interesting exporation of the work involved in building an ecosystem of open data, the challenges of making it sustainable, and the benefits of building for the long term rather than trying to achieve a quick win.
                                                                                        49 min
                                                                                      • Building The Open Data Ecosystem For Music And More At Metabrainz
                                                                                        Summary

                                                                                        The Musicbrainz project was an early entry in the movement to build an open data ecosystem. In recent years, the Metabrainz Foundation has fostered a growing ecosystem of projects to support the contribution of, and access to, metadata, listening habits, and review of music. The majority of those projects are written in Python, and in this episode Param Singh explains how they are built, how they fit together, and how they support the goals of the Metabrains Foundation. This was an interesting exporation of the work involved in building an ecosystem of open data, the challenges of making it sustainable, and the benefits of building for the long term rather than trying to achieve a quick win.

                                                                                        Announcements
                                                                                        • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
                                                                                        • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
                                                                                        • Before you put your code into production you need to make sure that it passes all of the tests, that it has been packaged with all of the dependencies, and that you haven’t introduced any security issues. Instead of running all of that on your laptop, let Codefresh handle it automatically with their continuous integration and continuous delivery platform. Built for the modern era of cloud-native computing, they make publishing to Kubernetes, serverless platforms, and virtual machines fast and seamless. With a growing library of pre-made steps, a flexible pipeline definition, and unlimited scale Codefresh lets you ship faster and safer than ever. Go to pythonpodcast.com/codefresh today to get unlimited builds on your free account.
                                                                                        • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
                                                                                        • Your host as usual is Tobias Macey and today I’m interviewing Param Singh about the ways that Python is being used across the various Metabrainz projects
                                                                                        • Interview
                                                                                          • Introductions
                                                                                          • How did you get introduced to Python?
                                                                                          • Can you start by giving an overview of what the Metabrainz organization is and the various projects that it encompasses?
                                                                                            • What are the motivations for creating those projects and some of the origin story for Metabrainz?
                                                                                            • The Musicbrainz server is the longest running project and is written in Perl. What was the reason for switching to Python for all of the other *brainz projects?
                                                                                            • How does the MetaBrainz Foundation sustain itself? Where do the funds come from?
                                                                                              • How do you determine where and how to allocate the funding that you receive?
                                                                                              • Which of the *brainz projects is the most complex or challenging to build, whether due to technical or sociological reasons?
                                                                                              • How do you source and manage the information that powers all of the Metabrainz projects?
                                                                                              • How is development of the various projects organized?
                                                                                                • How does that influence the amount of code sharing that is possible between them?
                                                                                                • Of the projects that you have been involved in, how are they architected?
                                                                                                  • What are the main ways that the projects differ in how they are implemented?
                                                                                                  • What are some of the ways that you are using Python in support of the various projects that you work on?
                                                                                                  • What are some of the most interesting, innovative, or unexpected ways that you have seen the projects or data built by Metabrainz being used?
                                                                                                  • What are some of the most interesting, unexpected, or challenging lessons that you have learned while working as a contributor and maintainer of the Metabrainz projects?
                                                                                                  • What is in store for the future of the existing Metabrainz projects?
                                                                                                  • What are the next domains that are being considered for building a Metabrainz platform for?
                                                                                                  • Keep In Touch
                                                                                                    • LinkedIn
                                                                                                    • paramsingh on GitHub
                                                                                                    • Website
                                                                                                    • Picks
                                                                                                      • Tobias
                                                                                                        • Beets music library organizer
                                                                                                          • Podcast Episode
                                                                                                          • Param
                                                                                                            • Prateek Kuhad
                                                                                                            • Links
                                                                                                              • Metabrainz
                                                                                                                • Musicbrainz
                                                                                                                • Listenbrainz
                                                                                                                • Acousticbrainz
                                                                                                                • Bookbrainz
                                                                                                                • Critiquebrainz
                                                                                                                • Picard
                                                                                                                • Stripe
                                                                                                                • The Himalayas
                                                                                                                • Dublin Ireland
                                                                                                                • XKCD Import Antigravity
                                                                                                                  • Antigravity Python Module
                                                                                                                  • Last.fm
                                                                                                                  • Google Summer of Code
                                                                                                                  • CDDB
                                                                                                                  • Perl
                                                                                                                  • Flask
                                                                                                                  • SQLAlchemy
                                                                                                                  • 3rd anniversary cake
                                                                                                                  • Redis
                                                                                                                  • PostgreSQL
                                                                                                                  • RabbitMQ
                                                                                                                  • Spark
                                                                                                                  • Music Technology Group
                                                                                                                  • Splunk
                                                                                                                  • Artist Origins Map on ListenBrainz
                                                                                                                  • The intro and outro music is from Requiem for a Fish The Freak Fandango Orchestra / CC BY-SA

                                                                                                                    49 min
                                                                                                                  • Growing Dask To Make Scaling Python Data Science Easier At Coiled
                                                                                                                    Python is a leading choice for data science due to the immense number of libraries and frameworks readily available to support it, but it is still difficult to scale. Dask is a framework designed to transparently run your data analysis across multiple CPU cores and multiple servers. Using Dask lifts a limitation for scaling your analytical workloads, but brings with it the complexity of server administration, deployment, and security. In this episode Matthew Rocklin and Hugo Bowne-Anderson discuss their recently formed company Coiled and how they are working to make use and maintenance of Dask in production. The share the goals for the business, their approach to building a profitable company based on open source, and the difficulties they face while growing a new team during a global pandemic.
                                                                                                                    53 min
                                                                                                                  • Growing Dask To Make Scaling Python Data Science Easier At Coiled
                                                                                                                    Summary

                                                                                                                    Python is a leading choice for data science due to the immense number of libraries and frameworks readily available to support it, but it is still difficult to scale. Dask is a framework designed to transparently run your data analysis across multiple CPU cores and multiple servers. Using Dask lifts a limitation for scaling your analytical workloads, but brings with it the complexity of server administration, deployment, and security. In this episode Matthew Rocklin and Hugo Bowne-Anderson discuss their recently formed company Coiled and how they are working to make use and maintenance of Dask in production. The share the goals for the business, their approach to building a profitable company based on open source, and the difficulties they face while growing a new team during a global pandemic.

                                                                                                                    Announcements
                                                                                                                    • Hello and welcome to Podcast.__init__, the podcast about Python and the people who make it great.
                                                                                                                    • When you’re ready to launch your next app or want to try a project you hear about on the show, you’ll need somewhere to deploy it, so take a look at our friends over at Linode. With the launch of their managed Kubernetes platform it’s easy to get started with the next generation of deployment and scaling, powered by the battle tested Linode platform, including simple pricing, node balancers, 40Gbit networking, dedicated CPU and GPU instances, and worldwide data centers. Go to pythonpodcast.com/linode and get a $60 credit to try out a Kubernetes cluster of your own. And don’t forget to thank them for their continued support of this show!
                                                                                                                    • This portion of Python Podcast is brought to you by Datadog. Do you have an app in production that is slower than you like? Is its performance all over the place (sometimes fast, sometimes slow)? Do you know why? With Datadog, you will. You can troubleshoot your app’s performance with Datadog’s end-to-end tracing and in one click correlate those Python traces with related logs and metrics. Use their detailed flame graphs to identify bottlenecks and latency in that app of yours. Start tracking the performance of your apps with a free trial at datadog.com/pythonpodcast. If you sign up for a trial and install the agent, Datadog will send you a free t-shirt.
                                                                                                                    • You listen to this show to learn and stay up to date with the ways that Python is being used, including the latest in machine learning and data analysis. For more opportunities to stay up to date, gain new skills, and learn from your peers there are a growing number of virtual events that you can attend from the comfort and safety of your home. Go to pythonpodcast.com/conferences to check out the upcoming events being offered by our partners and get registered today!
                                                                                                                    • Your host as usual is Tobias Macey and today I’m interviewing Matthew Rocklin and Hugo Bowne-Anderson about their work building a business around the Dask ecosystem at Coiled
                                                                                                                    • Interview
                                                                                                                      • Introductions
                                                                                                                      • How did you get introduced to Python?
                                                                                                                      • Can you give a quick overview of what Dask is and your motivations for creating it?
                                                                                                                        • How has Dask changed or evolved in the past 3 1/2 years since we last talked about it?
                                                                                                                        • How has the rest of the ecosystem changed in that time?
                                                                                                                        • After working on Dask for the past few years, what led you to the decision to build a business around it?
                                                                                                                        • What are the sharp edges of programming for Dask that users are looking for help on solving?
                                                                                                                        • What are the difficulties that users face in deploying and maintaining a production installation of Dask?
                                                                                                                        • What are the limitations of Dask when scaling both up and down?
                                                                                                                        • What are you building at Coiled to improve the user experience for users of Python and Dask?
                                                                                                                          • What are your thoughts on the pros and cons of orienting your messaging around the scalability of Python, as opposed to focusing on a specific industry or problem domain?
                                                                                                                          • What are the challenges that you are facing in managing the tensions between the open source and proprietary work that you are doing?
                                                                                                                          • How are you handling the ongoing governance of the Dask project?
                                                                                                                          • What are some of the most interesting, unexpected, or challenging lessons that you have learned while building and launching a company based on an open source project?
                                                                                                                          • What do you have planned for the future of both Coiled and Dask?
                                                                                                                          • Keep In Touch
                                                                                                                            • Matt
                                                                                                                              • Website
                                                                                                                              • @mrocklin on Twitter
                                                                                                                              • mrocklin on GitHub
                                                                                                                              • Hugo
                                                                                                                                • LinkedIn
                                                                                                                                • @hugobowne on Twitter
                                                                                                                                • Website
                                                                                                                                • Picks
                                                                                                                                  • Tobias
                                                                                                                                    • The Hobbit
                                                                                                                                      • Audiobook
                                                                                                                                      • Audible Free Trial (affiliate link)
                                                                                                                                      • Matt
                                                                                                                                        • Prefect
                                                                                                                                        • Hugo
                                                                                                                                          • Race After Technology by Ruha Benjamin
                                                                                                                                          • Ruha Benjamin on deep learning: Computational depth without sociological depth is ‘superficial learning’
                                                                                                                                          • Closing Announcements
                                                                                                                                            • Thank you for listening! Don’t forget to check out our other show, the Data Engineering Podcast for the latest on modern data management.
                                                                                                                                            • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                            • If you’ve learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                            • To help other people find the show please leave a review on iTunes and tell your friends and co-workers
                                                                                                                                            • Join the community in the new Zulip chat workspace at pythonpodcast.com/chat
                                                                                                                                            • Links
                                                                                                                                              • Sign up for the Coiled Beta!
                                                                                                                                              • Coiled
                                                                                                                                              • Dask
                                                                                                                                              • Data Engineering Podcast Interview About Dask
                                                                                                                                              • PyData
                                                                                                                                              • NumPy
                                                                                                                                              • SciPy
                                                                                                                                              • Cell Biology
                                                                                                                                              • Datacamp
                                                                                                                                              • Dataframed
                                                                                                                                              • Matthew Rocklin on Podcast.__init__ about functional programming with Toolz
                                                                                                                                              • IPython Notebook
                                                                                                                                              • PyTorch
                                                                                                                                                • Podcast Episode
                                                                                                                                                • Airflow
                                                                                                                                                • Prefect
                                                                                                                                                • XGBoost
                                                                                                                                                • Tornado
                                                                                                                                                • Coiled Blog Post About The Goals of Dask
                                                                                                                                                • Spark
                                                                                                                                                • AsyncIO
                                                                                                                                                • Concurrent.futures
                                                                                                                                                • Pangeo
                                                                                                                                                • Xarray
                                                                                                                                                • RAPIDS
                                                                                                                                                • Nvidia
                                                                                                                                                • Cuda
                                                                                                                                                • Prefect
                                                                                                                                                  • Data Engineering Podcast Episode
                                                                                                                                                  • Celery
                                                                                                                                                  • Life Sciences
                                                                                                                                                  • Tensorflow
                                                                                                                                                  • Snorkel
                                                                                                                                                    • Data Engineering Podcast Episode
                                                                                                                                                    • Dagster
                                                                                                                                                      • Data Engineering Podcast Episode
                                                                                                                                                      • DevOps
                                                                                                                                                      • Docker
                                                                                                                                                      • Kubernetes
                                                                                                                                                      • Metaflow
                                                                                                                                                        • Podcast Episode
                                                                                                                                                        • Ray
                                                                                                                                                          • Podcast Episode
                                                                                                                                                          • Anyscale
                                                                                                                                                          • Yarn
                                                                                                                                                          • Gartner Hype Cycle
                                                                                                                                                          • Travis Oliphant
                                                                                                                                                          • Postgres
                                                                                                                                                          • Amazon ECS
                                                                                                                                                          • Django
                                                                                                                                                          • Django Allauth
                                                                                                                                                          • Quansight
                                                                                                                                                          • Wes McKinney
                                                                                                                                                            • 53 min
                                                                                                                                                            • Supporting The Full Lifecycle Of Machine Learning Projects With Metaflow
                                                                                                                                                              Netflix uses machine learning to power every aspect of their business. To do this effectively they have had to build extensive expertise and tooling to support their engineers. In this episode Savin Goyal discusses the work that he and his team are doing on the open source machine learning operations platform Metaflow. He shares the inspiration for building an opinionated framework for the full lifecycle of machine learning projects, how it is implemented, and how they have designed it to be extensible to allow for easy adoption by users inside and outside of Netflix. This was a great conversation about the challenges of building machine learning projects and the work being done to make it more achievable.
                                                                                                                                                              45 min

                                                                                                                                                            About The Python Podcast.__init__

                                                                                                                                                            From the publisher's feed

                                                                                                                                                            The podcast about Python and the people who make it great

                                                                                                                                                            More shows like The Python Podcast.__init__

                                                                                                                                                            Freakonomics Radio by Freakonomics Radio + Stitcher

                                                                                                                                                            Freakonomics Radio

                                                                                                                                                            32,053 Listeners

                                                                                                                                                            Odd Lots by Bloomberg

                                                                                                                                                            Odd Lots

                                                                                                                                                            1,977 Listeners

                                                                                                                                                            The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                                            The Changelog: Software Development, Open Source

                                                                                                                                                            286 Listeners

                                                                                                                                                            Data Skeptic by Kyle Polich

                                                                                                                                                            Data Skeptic

                                                                                                                                                            476 Listeners

                                                                                                                                                            Software Engineering Daily by Software Engineering Daily

                                                                                                                                                            Software Engineering Daily

                                                                                                                                                            623 Listeners

                                                                                                                                                            Talk Python To Me by Michael Kennedy

                                                                                                                                                            Talk Python To Me

                                                                                                                                                            582 Listeners

                                                                                                                                                            Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                                            Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                                            305 Listeners

                                                                                                                                                            Python Bytes by Michael Kennedy and Calvin Hendryx-Parker

                                                                                                                                                            Python Bytes

                                                                                                                                                            213 Listeners

                                                                                                                                                            Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                                            Syntax - Tasty Web Development Treats

                                                                                                                                                            985 Listeners

                                                                                                                                                            DataFramed by DataCamp

                                                                                                                                                            DataFramed

                                                                                                                                                            265 Listeners

                                                                                                                                                            Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                                            Practical AI

                                                                                                                                                            202 Listeners

                                                                                                                                                            The Intelligence from The Economist by The Economist

                                                                                                                                                            The Intelligence from The Economist

                                                                                                                                                            2,543 Listeners

                                                                                                                                                            The Real Python Podcast by Real Python

                                                                                                                                                            The Real Python Podcast

                                                                                                                                                            139 Listeners

                                                                                                                                                            声动早咖啡 by 声动活泼

                                                                                                                                                            声动早咖啡

                                                                                                                                                            305 Listeners

                                                                                                                                                            The Foreign Affairs Interview by Foreign Affairs Magazine

                                                                                                                                                            The Foreign Affairs Interview

                                                                                                                                                            474 Listeners