O'Reilly Data Show Podcast

O'Reilly Data Show Podcast

By O'Reilly Media
Download on the App Store

O'Reilly Data Show Podcast episodes

  • Machine learning at Spotify: You are what you stream
    In this episode of the Data Show, I spoke with Christine Hung, head of data solutions at Spotify. Prior to joining Spotify, she led data teams at the NY Times and at Apple (iTunes). Having led teams at three different companies, I wanted to hear her thoughts on digital transformation, and I wanted to know how she approaches the challenge of building, managing, and nurturing data teams.
    I also wanted to learn more about what goes into building a recommender system for a popular consumer service like Spotify. Engagement should clearly be the most important metric, but there are other considerations, such as introducing users to new or “long tail” content.
    Here are some highlights from our conversation:
    Recommenders at Spotify
    For us, engagement always comes first. At Spotify, we have a couple hundred people who are just focused on user engagement, and this is the group that creates personalized playlists, like Discover Weekly or your Daily Mix for you. We know our users love discovery and see Spotify as a very important platform for them to discover something new, but there are also times when people just want to have some music played in the background that fits the mood. But again, we don’t have a specific agenda in terms of what we should push for. We want to give you what you want so that you are happy, which is why we invested so much in understanding people through music. If we believe you might like some “long tail” content, we will recommend it to you because it makes you happy, but we can also do the same for the top 100 track if we believe you will enjoy them.
    Music is like a mirror
    Music is like a mirror, and it tells people a lot about who you are and what you care about, whether you like it or not. We love to say “you are what you stream,” and that is so true. As you can imagine, we invest a lot in our machine learning capabilities to predict people’s preference and context, and of course, all the data we use to train the model is anonymized. We take in large amounts of anonymized training data to develop these models, and we test them out with different uses cases, analyze results, and use the learning to improve those models.
    Just to give you my personal example to illustrate how it works, you can learn a lot about me just by me telling you what I stream. You will see that I use my running playlist only during the weekend in early mornings, and I have a lot of children’s songs streamed at my house between 5 p.m. and 7 p.m. I also have a lot of tango and salsa playlists that I created and followed. So what does that tell you? It tells you that I am probably a weekend runner, which means I have some kind of affiliation for fitness; it tells you that I am probably a mother and play songs for my child after I get home from work; it also tells you that I somehow like tango and salsa, so I am probably a dancer, too. As you can see, we are investing a lot into understanding people’s context and preference so we can start capturing different moments of their lives. And, of course, the more we understand your context, your preference, and what you are looking for, the better we can customize your playlists for you.
    Related resources:
    Music, the window into your soul: Christine Hung’s keynote at Strata Data NYC 2017
    “Transforming organizations through analytics centers of excellence”: Carme Artigas on helping enterprises transform themselves with big data tools and technologies.
    “A framework for building and evaluating data products”: Grace Huang on lessons learned in the course of machine learning product launches.
    “How companies can navigate the age of machine learning”: to become a “machine learning company,” you need tools and processes to overcome challenges in data, engineering, and models.
    22 min
  • The current state of Apache Kafka
    In this episode of the Data Show, I spoke with Neha Narkhede, co-founder and CTO of Confluent. As I noted in a recent post on “the age of machine learning,” data integration and data enrichment are non-trivial and ongoing challenges for most companies. Getting data ready for analytics—including machine learning—remains an area of focus for most companies. It turns out, “data lakes” have become staging grounds for data; more refinement usually needs to be done before data is ready for analytics. By making it easier to create and productionize data refinement pipelines on both batch and streaming data sources, analysts and data scientists can focus on analytics that can unlock value from data.
    On the open source side, Apache Kafka continues to be a popular framework for data ingestion and integration. Narkhede was part of the team that created Kafka, and I wanted to get her thoughts on where this popular framework is headed.
    Here are some highlights from our conversation:
    The first engineering project that made use of Apache Kafka
    If I remember correctly, we were putting Hadoop into a place at LinkedIn for the first time, and I was on the team that was responsible for that. The problem was that all our scripts were actually built for another data warehousing solution. The questions was, are we going to rewrite all of those scripts and now sort of make them Hadoop specific? And what happens when a third and a fourth and a fifth system is put into place?
    So, the initial motivating use case was: ‘we are putting this Hadoop thing into place. That’s the new-age data warehousing solution. It needs access to the same data that is coming from all our applications. So, that is the thing we need to put into practice.’ This became Kafka’s very first use case at LinkedIn. From there, because that was very easy and I actually helped move one of the very first workloads to Kafka, it was hardly difficult to convince the rest of the LinkedIn engineering team to start moving over to Kafka.
    So from there, Kafka adoption became pretty vital. Now, I think years down the line, all of LinkedIn runs on Kafka. It’s essentially the central nervous system for the whole company.
    Microservices and Kafka
    My own opinion of microservices is that it lets you add more money and turn it into software at a more constant rate by allowing engineers to focus on various parts of the application, by essentially decoupling a big monolith so that a lot of things can happen in parallel development of real applications.
    … The upside is that it lets you move fast. It adds a certain amount of agility to an engineering organization. But it comes with its own set of challenges. And these were not very obvious back then. How are all these microservices deployed? How are they monitored? And, most importantly, how do they communicate with each other? The communication bit is where Kafka comes in. When you break a monolith, you break state. And you distribute that state across different machines that run all those different applications.
    So now the problem is, ‘well, how do these microservices share that state? How do they talk to each other?’ Frequently, the expectation is that things happens in real time. The context of microservices where streams or Kafka comes in is in the communication model for those microservices. I should just say that there isn’t a one size fits all when it comes to communication patterns for microservices.
    Related resources:
    Kafka: The Definitive Guide
    “Architecting and building end-to-end streaming applications“: Karthik Ramasamy on Heron, DistributedLog, and designing real-time applications.
    “Semi-supervised, unsupervised, and adaptive algorithms for large-scale time series“: Ira Cohen on developing machine learning tools for a broad range of real-time applications.
    “Building Apache Kafka from scratch“: Jay Kreps on data integration, event data, and the Internet of Things.
    I Logs
    38 min
  • Building a natural language processing library for Apache Spark
    When I first discovered and started using Apache Spark, a majority of the use cases I used it for involved unstructured text. The absence of libraries meant rolling my own NLP utilities, and, in many cases, implementing a machine learning library (this was pre deep learning, and MLlib was much smaller). I’d always wondered why no one bothered to create an NLP library for Spark when many people were using Spark to process large amounts of text. The recent, early success of BigDL confirms that users like the option of having native libraries.
    In this episode of the Data Show, I spoke with David Talby of Pacific.AI, a consulting company that specializes in data science, analytics, and big data. A couple of years ago I mentioned the need for an NLP library within Spark to Talby; he not only agreed, he rounded up collaborators to build such a library. They eventually carved out time to build the newly released Spark NLP library. Judging by the reception received by BigDL and the number of Spark users faced with large-scale text processing tasks, I suspect Spark NLP will be a standard tool among Spark users.
    Talby and I also discussed his work helping companies build, deploy, and monitor machine learning models. Tools and best practices for model development and deployment are just beginning to emerge—I summarized some of them in a recent post, and, in this episode, I discussed these topics with a leading practitioner.
    Here are some highlights from our conversation:
    The state of NLP in Spark
    Here are your two choices today. Either you want to leverage all of the performance and optimization that Spark gives you, which means you want to stay basically within the JVM, and you want to use a Java-based library. In which case, you have options that include OpenNLP, which is open source, or Stanford NLP, which requires licensing in order to use in a commercial product. These are older and more academically oriented libraries. So, they have limitations in performance and what they do.
    Another option is to look at something like spaCy—a Python-based library that really has raised the bar in terms of usability, and the trade-offs between analytical accuracy and performance. But then your challenge is that you have your text in Spark, but to call the spaCy pipeline, you basically have to move the data from the JVM to a Python process, do some processing there, and send it back, which in practice means you take a huge performance hit because most of the processing you do is really moving strings between operating system processors.
    … So, really what we were looking for is a solution to work on text directly, within a data frame. A tool that will take into account everything Spark gives in terms of caching, distributed computation, and the other optimizations. This enable users to basically run an NLP and machine learning pipeline directly on their text.
    Enter Spark NLP
    Spark NLP. Image by David Talby, used with permission.
    The core purpose of an NLP library is the ability to take text and then apply a set of annotations on the text. So, the basic annotations we ship in this initial version of Spark NLP include things like a tokenizer, a lemmatizer, sentence boundary detection, and paragraph boundary detection. Then on top of that, we include things like sentiment analysis, spell checker so we can auto-suggest corrections, and a dependency parser so we can not just know that we have a noun and a verb, but also that this verb talks about the specific noun, which is often semantically interesting. We also include named entity recognition algorithms.
    Deploying and monitoring machine learning models in production
    I think what’s happening is that people expect basic model development to be very similar to software development. When we started doing software development, we started it wrong. We assumed software engineering was a lot like civil engineering or mechanical engineering. It took a good 30 years until we said no, this is actual
    34 min
  • Machine intelligence for content distribution, logistics, smarter cities, and more
    In this episode of the Data Show, I spoke with Rhea Liu, analyst at China Tech Insights, a new research firm that is part of Tencent’s Online Media Group. If there’s one place where AI and machine learning are discussed even more than the San Francisco Bay Area, that would be China. Each time I go to China, there are new applications that weren’t widely available just the year before. This year, it was impossible to miss bike sharing, mobile payments seemed to be accepted everywhere, and people kept pointing out nascent applications of computer vision (facial recognition) to identity management and retail (unmanned stores).
    I wanted to consult local market researchers to help make sense of some of the things I’ve been observing from afar. Liu and her colleagues have put out a series of interesting reports highlighting some of these important trends. They also have an annual report—Trends & Predictions for China’s Tech Industry in 2018—that Liu will discuss in her keynote and talk at Strata Data Singapore in December.
    Here are some highlights from our conversation:
    Machine learning and content distribution
    Media consumption takes a large proportion of people’s everyday life here in China. Before, people learned their news from news portals and from editorial teams who served as the gatekeepers. People now trust machine learning algorithms with editorial and agenda setting. Apps like Toutiao have become very popular.
    It’s been quite a surprise to most news portals and media professionals here in China. People are trying to find a balance between the traditional ways of content creation and the new ways of content distribution by aggregators fully powered by machines. Toutiao’s news recommendation engine is purely a black box to most people. … But users are spending more and more time on these types of platforms. And, machine-generated news feeds have become a big thing.
    … So, it’s now becoming a content war again. After these algorithms improve the efficiency of content distribution, the battle may come down to what content you have.
    Bike sharing
    Bike sharing is kind of a new model adapted to Chinese society. … In between every subway station, there’s still several miles to go, where people still need to walk or maybe take a taxi. Bike sharing is being used to replace these other kinds of approaches.
    There are two primary players. One is Mobike and the other one is Ofo, and they started with different models, actually. Ofo started a year or two earlier from a university campus. … It provided this kind of public bike rental system to users on campus. This was kind of the preliminary prototype of this model. Mobike started in a city.
    These bike sharing companies have their GPS systems on the bikes, and the bikes have digital electronic locks that can be unlocked with an app on your phone. These technologies, combined together, can help them collect data as well as have a better management system of all the bikes they distribute over a city.
    Smart cities
    It’s still a maybe, but it’s very likely we are going to include things about smart cities in our 2018 reports. … This includes AR applications to help build better cities for urban planning. … Urban planning is a very complicated thing, and what we are missing there is, we can be a little bit left behind because of the lack of data. But now people have different types of data. For example, I know the ride sharing company Didi is collaborating with several city governments to help them do urban planning: by using data to better understand traffic, how to manage traffic light systems in the city, and also the bus system.
    City governments at all levels are now collaborating with all these tech companies to explore applications of their data to improve the cities we have in China. … This is going to be a very important opportunity for the tech companies here in China, especially in terms of their data applica
    37 min
  • Vehicle-to-vehicle communication networks can help fuel smart cities
    In this episode of the Data Show, I spoke with Bruno Fernandez-Ruiz, co-founder and CTO of Nexar. We first met when he was leading Yahoo! technical teams charged with delivering a variety of large-scale, real-time data products. His new company is helping build out critical infrastructure for the emerging transportation sector.
    While some question whether V2X communication is necessary to get to fully autonomous vehicles, Nexar is already paving the way by demonstrating how a vehicle-to-vehicle (V2V) communication network can be built efficiently. As Fernandez-Ruiz points out, there are many applications for such a V2V network (safety being the most obvious one). I’m particularly fascinated by what such a network, and the accompanying data, opens up for future, smarter cities. As I pointed out in a post on continuous learning, simulations are an important component of training AI applications. It seems reasonable to expect that the data sets collected by V2V networks will be useful for smart city planners of the future.
    Here are some highlights from our conversation:
    The many applications of a vehicle-to-vehicle network
    Imagine if every vehicle on the road was equipped with a transponder that allowed it to connect to a network and say: ‘Hey, here I am. This is where I’m going. This is how fast I’ve gone. This is where I’ve been in the last 10 seconds. This is where I think I’m going to be in the next 10 seconds.’ Now imagine you were sharing that with all the vehicles around you so all these vehicles can predict, react, and even proact how they should behave on the road.
    You can start solving for safety and for traffic congestion. You start solving for utilization of the road, for pollution, and many other problems. … That’s kind of the vision. I think with autonomous vehicles on the road, this will be even more important. You’ll have humans sharing the road with them—then, how does this mix of human and autonomous vehicles talk to each other?
    … So, we’re trying to just build that vehicle-to-vehicle network. You can call it ‘ground traffic control.’ It’s similar to what happened in the air when radar and beaconing technology became available, and people said, ‘Well, we should probably connect to these things to be able to know where the planes are and tell them where to go.’
    Redundancy and safety
    There are many situations in which other autonomous technologies actually may fail, and I think that’s where vehicle-to-vehicle communication becomes both a necessary technology for the true future of complete automation. It’s also a redundant system for when that camera, or for when that radar, when whatever other sensors in the car fail.
    Related resources:
    “Creating autonomous vehicle systems“: Understanding AV technologies and how to integrate them
    Cars that coordinate with people: 2017 AI Conference keynote by Anca Dragan
    “How big data and AI will reshape the automotive industry“
    “How intelligent data platforms are powering smart cities“
    “Why continuous learning is key to AI“: A look ahead at the tools and methods for learning from sparse feedback.
    46 min
  • Transforming organizations through analytics centers of excellence
    In this episode of the Data Show, I spoke with Carme Artigas, co-founder and CEO of Synergic Partners (a Telefonica company). As more companies adopt big data technologies and techniques, it’s useful to remember that the end goal is to extract information and insight. In fact, as with any collection of tools and technologies, the main challenge is identifying and prioritizing use cases.
    As Artigas describes, one can categorize use cases for big data into the following types:
    Improve decision-making or operational efficiency
    Generate new or additional revenue
    Predict or prevent fraud (forecasting or minimizing risks)
    Artigas has spent many years helping large organizations develop best practices for how to use data and analytics. We discussed some of the key challenges faced by organizations that wish to adopt big data technologies, centers of excellence for analytics, and AI in the enterprise.
    Here are some highlights from our conversation:
    Adopting big data analytics: Remaining key challenges
    For me, the first challenge is that there’s a lack of skills across organizations. We know there’s a global shortage of analytic talent, so it’s not only challenging for a company to acquire the right talent, but also to make sure that this talent is accessible across the organization. It’s usually concentrated in some departments, and it’s very difficult to leverage those skills for the good of the entire organization.
    The second challenge I see is lack of standards and lack of governance. You might find that every single data science team uses their own libraries or their own version of code or their own software tools. They are thinking about the benefit for a particular use case, and the best tools and the best models for that particular use case. But this compartmentalized approach cannot scale up; having a variety of versions of libraries and tools make it very, very difficult to industrialize and implement global solutions.
    Finally, new skills you need to develop are not only on the technical side—they are mostly on the business side. Decision-makers need to make decisions in different ways. They need to make decisions based on data, based on facts.
    Center of excellence for analytics
    The analytics center of excellence is a team of business and technical people that can be internal, external, and even crowd sourced. They have some centralized capabilities and also some distributed capabilities and resources, creating a common (online) workspace where they share methodologies, tools, models, and techniques. The objective is to gain efficiency and be able to implement initiatives across to the different business units. We have two main components of these centers of excellence: the business transformation unit and the deployment units.
    The business transformation unit (BTU) is the primary link of the center of excellence with the underlying business. We create ambassadors, and these ambassadors, who are part of the BTU, are responsible for identifying and prioritizing all business use cases. Then they connect with the deployment units—which we call cells—and these cells can grow organically during a project. So, first of all, the center of excellence must be connected with business. … We also create a centralized function called the ‘core team’ and an expansion unit called the ‘extended team.’ We have a few types of cells: the analytical cells, the operational cells, and the data utilization cells. So, it’s a way of concentrating the resources, gaining operational efficiency, having a center of know-how transferred to the rest of the organization, and ensuring best practices and methodologies.
    … A center of excellence is not a physical place. The center of excellence is a network of people who can be distributed in different geographies.
    Related resources:
    “The stages of enterprise IoT adoption“: Teresa Tung on building a business case for the Internet of Things
    “Data go
    39 min
  • The state of machine learning in Apache Spark
    In this episode of the Data Show, we look back to a recent conversation I had at the Spark Summit in San Francisco with Ion Stoica (UC Berkeley professor and executive chairman of Databricks) and Matei Zaharia (assistant professor at Stanford and chief technologist of Databricks). Stoica and Zaharia were core members of UC Berkeley’s AMPLab, which originated Apache Spark, Apache Mesos, and Alluxio.
    We began our conversation by discussing recent academic research that would be of interest to the Apache Spark community (Stoica leads the RISE Lab at UC Berkeley, Zaharia is part of Stanford’s DAWN Project). The bulk of our conversation centered around machine learning. Like many in the audience, I was first attracted to Spark because it simultaneously allowed me to scale machine learning algorithms to large data sets while providing reasonable latency.
    Here is a partial list of the items we discussed:
    The current state of machine learning in Spark.
    Given that a lot of innovation has taken place outside the Spark community (e.g., scikit-learn, TensorFlow, XGBoost), we discussed the role of Spark ML moving forward.
    The plan to make it easier to integrate advanced analytics libraries that aren’t “textbook machine learning,” like NLP, time series analysis, and graph analysis into Spark and Spark ML pipelines.
    Some upcoming projects from Berkeley and Stanford that target AI applications (including newer systems that provide lower latency, higher throughput).
    Recent Berkeley and Stanford projects that address two key bottlenecks in machine learning—lack of training data, and deploying and monitoring models in production.
    [Full disclosure: I am an advisor to Databricks.]
    Related resources:
    Spark: The Definitive Guide
    Advanced Analytics with Spark
    High-performance Spark
    Learning Path: Get Started with Natural Language Processing Using Python, Spark, and Scala
    Learning Path: Getting Up and Running with Apache Spark
    “The current state of applied data science”
    22 min
  • Effective mechanisms for searching the space of machine learning algorithms
    In this episode of the Data Show, I spoke with Ken Stanley, founding member of Uber AI Labs and associate professor at the University of Central Florida. Stanley is an AI researcher and a leading pioneer in the field of neuroevolution—a method for evolving and learning neural networks through evolutionary algorithms. In a recent survey article, Stanley went through the history of neuroevolution and listed recent developments, including its applications to reinforcement learning problems.
    Stanley is also the co-author of a book entitled Why Greatness Cannot Be Planned: The Myth of the Objective—a book I’ve been recommending to anyone interested in innovation, public policy, and management. Inspired by Stanley’s research in neuroevolution (into topics like novelty search and open endedness), the book is filled with examples of how notions first uncovered in the field of AI can be applied to many other disciplines and domains.
    The book closes with a case study that hits closer to home—the current state of research in AI. One can think of machine learning and AI as a search for ever better algorithms and models. Stanley points out that gatekeepers (editors of research journals, conference organizers, and others) impose two objectives that researchers must meet before their work gets accepted or disseminated: (1) empirical: their work should beat incumbent methods on some benchmark task, and (2) theoretical: proposed new algorithms are better if they can be proven to have desirable properties. Stanley argues this means that interesting work (“stepping stones”) that fail to meet either of these criteria fall by the wayside, preventing other researchers from building on potentially interesting but incomplete ideas.
    Here are some highlights from our conversation:
    Neuroevolution today
    In the state of the art today, the algorithms have the ability to evolve variable topologies or different architectures. There are pretty sophisticated algorithms for evolving the architecture of a neural network; in other words, what’s connected to what, not just what the weight of those connections are—which is what deep learning is usually concerned with.
    There’s also an idea of how to encode very, very large patterns of connectivity. This is something that’s been developed independently in neuroevolution where there’s not a really analogous thing in deep learning right now. This is the idea that if you’re evolving something that’s really large, then you probably can’t afford to encode the whole thing in the DNA. In other words, if we have 100 trillion connections in our brains, our DNA does not have 100 trillion genes. In fact, it couldn’t have a 100 trillion genes. It just wouldn’t fit. That would be astronomically too high. So then, how is it that with a much, much smaller space of DNA, which is about 30,000 genes or so, three billion base pairs, how would you get enough information in there to encode something that’s 100 trillion parts?
    This is the issue of encoding. We’ve become sophisticated at creating artificial encodings that are basically compressed in an analogous way, where you can have a relatively short string of information to describe a very large structure that comes out—in this case, a neural network. We’ve gotten good at doing encoding and we’ve gotten good at searching more intelligently through the space of possible neural networks. We originally thought what you need to do is just breed by choosing among the best. So, you say, ‘Well, there’s some task we’re trying to do and I’ll choose among the best to create the next generation.’
    We’ve learned since then that that’s actually not always a good policy. Sometimes you want to explicitly choose for diversity. In fact, that can lead to better outcomes.
    The myth of the objective
    Our book does recognize that sometimes pursuing objectives is a rational thing to do. But I think the
    46 min
  • How Ray makes continuous learning accessible and easy to scale
    In this episode of the Data Show, I spoke with Robert Nishihara and Philipp Moritz, graduate students at UC Berkeley and members of RISE Lab. I wanted to get an update on Ray, an open source distributed execution framework that makes it easy for machine learning engineers and data scientists to scale reinforcement learning and other related continuous learning algorithms. Many AI applications involve an agent (for example a robot or a self-driving car) interacting with an environment. In such a scenario, an agent will need to continuously learn the right course of action to take for a specific state of the environment.
    What do you need in order to build large-scale continuous learning applications? You need a framework with low-latency response times, one that is able to run massive numbers of simulations quickly (agents need to be able explore states within an environment), and supports heterogeneous computation graphs. Ray is a new execution framework written in C++ that contains these key ingredients. In addition, Ray is accessible via Python (and Jupyter Notebooks), and comes with many of the standard reinforcement learning and related continuous learning algorithms that users can easily call.
    As Nishihara and Moritz point out, frameworks like Ray are also useful for common applications such as dialog systems, text mining, and machine translation. Here are some highlights from our conversation:
    Tools for reinforcement learning
    Ray is something we’ve been building that’s motivated by our own research in machine learning and reinforcement learning. If you look at what researchers who are interested in reinforcement learning are doing, they’re largely ignoring the existing systems out there and building their own custom frameworks or custom systems for every new application that they work on.
    … For reinforcement learning, you need to be able to share data very efficiently, without copying it between multiple processes on the same machine, you need to be able to avoid expensive serialization and deserialization, and you need to be able to create a task and get the result back in milliseconds instead of hundreds of milliseconds. So, there are a lot of little details that come up.
    … In fact, people often use MPI along with lower-level multi-processing libraries to build the communication infrastructure for their reinforcement learning applications.
    Scaling machine learning in dynamic environments
    I think right now when we think of machine learning, we often think of supervised learning. But a lot of machine learning applications are changing from making just one prediction to making sequences of decisions and taking sequences of actions in dynamic environments.
    The thing that’s special about reinforcement learning is it’s not just the different algorithms that are being used, but rather the different problem domain that it’s being applied to: interactive, dynamic, real-time settings bring up a lot of new challenges.
    … The set of algorithms actually goes even a little bit further. Some of these techniques are even useful in, for example, things like text summarization and translation. You can use these techniques that have been developed in the context of reinforcement learning to better tackle some of these more classical problems [where you have some objective function that may not be easily differentiable].
    … Some of the classic applications that we have in mind when we think about reinforcement learning are things like dialogue systems, where the agent is one participant in the conversation. Or robotic control, where the agent is the robot itself and it’s trying to learn how to control its motion.
    … For example, we implemented the evolution algorithm described in a recent OpenAI paper in Ray. It was very easy to port to Ray, and writing it only took a couple of hours. Then we had a distributed implementation that scaled very well and we ran it on up to 1
    19 min
  • Why AI and machine learning researchers are beginning to embrace PyTorch
    In this episode of the Data Show, I spoke with Soumith Chintala, AI research engineer at Facebook. Among his many research projects, Chintala was part of the team behind DCGAN (Deep Convolutional Generative Adversarial Networks), a widely cited paper that introduced a set of neural network architectures for unsupervised learning. Our conversation centered around PyTorch, the successor to the popular Torch scientific computing framework. PyTorch is a relatively new deep learning framework that is fast becoming popular among researchers. Like Chainer, PyTorch supports dynamic computation graphs, a feature that makes it attractive to researchers and engineers who work with text and time-series.
    Here are some highlights from our conversation:
    The origins of PyTorch
    TensorFlow addressed one part of the problem, which is quality control and packaging. It offered a Theano style programming model, so it was a very low-level deep learning framework. … There are a multitude of front ends that are trying to cope with the fact that TensorFlow is a very low-level framework—there’s TF-slim, there’s Keras. I think there’s like 10 or 15, and just from Google there’s probably like four or five of those.
    On the Torch side, the philosophy has always been slightly different than Theano. I see TensorFlow as a much better Theano-style framework, and on the Torch side we had a philosophy that we want to be imperative, which means that you run your computation immediately. Debugging should be butter smooth. The user should never have trouble debugging their programs, whether they use a Python debugger or something like the GDB or something else.
    … Chainer was a huge inspiration. PyTorch is inspired primarily by three frameworks. Within the Torch community, certain researchers from Twitter built an auxiliary package called Autograd, and this was actually based on a package called Autograd in the Python community. Like Chainer, Autograd and Torch Autograd, all used a certain technique called tape-based automatic differentiation: that is, you have a tape recorder that records what operations you have performed and then it replays it backward to compute your gradients. This is a technique that is not used by any of the other major frameworks except PyTorch and Chainer. All of the other frameworks use what we call a static graph—that is, the user builds a graph, then they give that graph to an execution engine that is provided by the framework, and the framework executes it. It can analyze it ahead of time.
    These are very two different techniques. The tape-based differentiation gives you easier debuggability, and it gives you certain things that are more powerful (e.g., dynamic neural networks). The static graph-based approach gives you easier deployment to mobile, easier deployment to more exotic architectures, the ability to do compiler techniques ahead of time, and so on.
    Deep learning frameworks within Facebook
    Internally at Facebook, we have a unified strategy. We say PyTorch is used for all of research and Caffe 2 is used for all of production. This makes it easier for us to separate out which team does what and which tools do what. What we are seeing is, users first create a PyTorch model. When they are ready to deploy their model into production, they just convert it into a Caffe 2 model, then ship into either mobile or another platform.
    PyTorch user profiles
    PyTorch has gotten its biggest adoption from researchers, and it’s gotten about a moderate response from data scientists. As we expected, we did not get any adoption from product builders because PyTorch models are not easy to ship into mobile, for example. We also have people who we did not expect to come on board, like folks from OpenAI and several universities.
    Related resources:
    Building intelligent applications with deep learning and TensorFlow
    BigDL: Deep learning for Apache Spark
    MXNet: Deep learning that’s easy to implement and easy to
    37 min

About O'Reilly Data Show Podcast

From the publisher's feed

The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.

More shows like O'Reilly Data Show Podcast

Data Skeptic by Kyle Polich

Data Skeptic

476 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

624 Listeners

O'Reilly Radar Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Radar Podcast - O'Reilly Media Podcast

35 Listeners

O'Reilly Design Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Design Podcast - O'Reilly Media Podcast

8 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Machine Learning Guide by OCDevel

Machine Learning Guide

774 Listeners

DataFramed by DataCamp

DataFramed

265 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

202 Listeners

AWS Podcast by Amazon Web Services

AWS Podcast

202 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

203 Listeners

Last Week in AI by Skynet Today

Last Week in AI

316 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

99 Listeners

MIT Technology Review Narrated by MIT Technology Review

MIT Technology Review Narrated

262 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

681 Listeners

Practical News: AI & Business News by Practical News

Practical News: AI & Business News

25 Listeners