O'Reilly Data Show Podcast

O'Reilly Data Show Podcast

By O'Reilly Media
Download on the App Store

O'Reilly Data Show Podcast episodes

  • How big data and AI will reshape the automotive industry
    In this episode of the Data Show, I spoke with Evangelos Simoudis, co-founder of Synapse Partners and a frequent contributor to O’Reilly. He recently published a book entitled The Big Data Opportunity in Our Driverless Future, and I wanted get his thoughts on the transportation industry and the role of big data and analytics in its future. Simoudis is an entrepreneur, and he also advises and invests in many technology startups. He became interested in the automotive industry long before the current wave of autonomous vehicle startups was in the planning stages.
    Here are some highlights from our conversation:
    Understanding the automotive industry
    The more I started spending time with the automotive industry, the more I came to realize that, because of the autonomous vehicle technology and because of various forms of mobility services, which are stemming from new business models, the incumbent automotive industry is in significant risk of being disrupted.
    If you were to look at the automotive industry, the first thing that is very striking is that there’s a small number of very large companies that control a number of different labels. With GM, we talk about Chevy, we talk about Buick, we talk about Opel in Europe. There are a very small number of companies that control this trillion dollar industry.
    The other thing that is interesting is that these companies are responsible for designing the vehicle, manufacturing it, assembling it, post-manufacturing, and then creating demand, whereas the sale of the vehicle is done through the dealers. And they’re paying relatively little attention to what happens post-sale. So, that means there is a relatively little understanding of consumer behavior.
    The third observation is that the reason there are so few of these companies is because starting one is very capital intensive. And if you look at how much money, for example, a company like Tesla has been able to raise, you get a sense of what kind of capital is necessary. And the next point is that even though there is a lot of capital that’s being raised, in the end this is a relatively low margin business. Where you try to make it up is in volume. That’s why, if you look at all these corporations, they have extremely sophisticated supply chains, extremely sophisticated manufacturing lines, highly optimized, because they are working on maintaining these margins.
    Infrastructure for autonomous vehicles
    A vehicle needs to know very much what’s happening around it. So that means it needs to receive signals from roads, bridges, other vehicles. … The term people use is V2X or vehicle-to-everything communication.
    It will take a very long time to have the preponderance of vehicles being autonomous. So, we need infrastructure that will enable cars to safely operate in a hybrid world between autonomous vehicles and manually operated vehicles. I think the experiments that today involve just a few tens of cars will expand over the next few years. And I think the result of those experiments will give us an understanding and appreciation of the investments that we need to make and how to prioritize them, as well as the regulations that we will need to institute in order to have this type of hybrid environment operate safely.
    AI and big data
    The argument that I’m making, and this actually comes from my education on AI and my work on AI since the mid ‘80s, is that while machine learning is important, I think everybody needs to appreciate that it’s not only about machine learning. In order to bring to realization an autonomous vehicle, you need more than machine learning. And, of course, within machine learning we have neural network learning and particularly deep learning, and these are very important areas.
    But people need to realize that an autonomous vehicle requires the ability to plan, requires the ability to reason, to represent knowledge, to search. All of these are components of AI. What I’m ho
    52 min
  • A framework for building and evaluating data products

    In this episode of the Data Show, I spoke with Grace Huang, data science lead at Pinterest. With its combination of a large social graph, enthusiastic users, and multimedia data, I’ve long regarded Pinterest as a fascinating lab for data science. Huang described the challenge of building a sustainable content ecosystem and shared lessons from the front lines of machine learning product launches. We also discussed recommenders, the emergence of deep learning as a technique used within Pinterest, and the role of data science within the company.

    23 min
  • Building a next-generation platform for deep learning

    In this episode of the Data Show, I speak with Naveen Rao, VP and GM of the Artificial Intelligence Products Group at Intel. In an earlier episode, we learned that scaling current deep learning models requires innovations in both software and hardware. Through his startup Nervana (since acquired by Intel), Rao has been at the forefront of building a next generation platform for deep learning and AI.

    I wanted to get his thoughts on what the future infrastructure for machine learning would look like. At least for now, we’re seeing a variety of approaches, and many companies are using heterogeneous processors (even specialized ones) and proprietary interconnects for deep learning. Nvidia and Intel Nervana are set to release processors that excel at both training and inference, but as Rao pointed out, at large-scale there are many considerations—including utilization, power consumption, and convenience—that come into play.

    28 min
  • A scalable time-series database that supports SQL

    In this episode of the Data Show, I spoke with Michael Freedman, CTO of Timescale and professor of computer science at Princeton University. When I first heard that Freedman and his collaborators were building a time-series database, my immediate reaction was: “Don’t we have enough options already?” The early incarnation of Timescale was a startup focused on IoT, and it was while building tools for the IoT problem space that Freedman and the rest of the Timescale team came to realize that the database they needed wasn’t available (at least out in open source). Specifically, they wanted a database that could easily support complex queries and the sort of real-time applications many have come to associate with streaming platforms. Based on early reactions to TimescaleDB, many users concur.

    50 min
  • Programming collective intelligence for financial trading

    In this episode of the Data Show, I spoke with Geoffrey Bradway, VP of engineering at Numerai, a new hedge fund that relies on contributions of external data scientists. The company hosts regular competitions where data scientists submit machine learning models for classification tasks. The most promising submissions are then added to an ensemble of models that the company uses to trade in real-world financial markets.

    27 min
  • Creating large training data sets quickly

    In this episode of the Data Show, I spoke with Alex Ratner, a graduate student at Stanford and a member of Christopher Ré’s Hazy research group. Training data has always been important in building machine learning algorithms, and the rise of data-hungry deep learning models has heightened the need for labeled data sets. In fact, the challenge of creating training data is ongoing for many companies; specific applications change over time, and what were gold standard data sets may no longer apply to changing situations.

    48 min
  • Data science and deep learning in retail

    In this episode of the Data Show, I spoke with Jeremy Stanley, VP of data science at Instacart, a popular grocery delivery service that is expanding rapidly. As Stanley describes it, Instacart operates a four-sided marketplace comprised of retail stores, products within the stores, shoppers assigned to the stores, and customers who order from Instacart. The objective is to get fresh groceries from popular retailers delivered to customers in a timely fashion. Instacart’s goals land them in the center of the many opportunities and challenges involved in building high-impact data products.

    50 min
  • Language understanding remains one of AI’s grand challenges

    In this episode of the Data Show, I spoke with David Ferrucci, founder of Elemental Cognition and senior technologist at Bridgewater Associates. Ferrucci served as principal investigator of IBM’s DeepQA project and led the Watson team that became champion of the Jeopardy! quiz show. Elemental Cognition (EC) is a research group focused on building an AI system that will be equipped with state-of-the-art natural language understanding technologies. Ferrucci envisions that EC will ship with foundational knowledge in many subject areas, but will be able to very quickly acquire knowledge in other (specialized) domains with the help of “human mentors.”

    Having built and deployed several prominent AI systems through the years, I also wanted to get Ferrucci’s perspective on the evolution of AI technologies, and how enterprises can take advantage of all the exciting recent developments.

    39 min
  • Data preparation in the age of deep learning

    In this episode of the Data Show, I spoke with Lukas Biewald, co-founder and chief data scientist at CrowdFlower. In a previous episode we covered how the rise of deep learning is fueling the need for large labeled data sets and high-performance computing systems. CrowdFlower has a service that many leading companies have come to rely on to provide them with labeled data sets to train machine learning models. As deep learning models get larger and more complex, they require training data sets that are bigger than those required by other machine learning techniques.

    37 min
  • Scaling machine learning

    In this episode of the Data Show, I spoke with Reza Zadeh, adjunct professor at Stanford University, co-organizer of ScaledML, and co-founder of Matroid, a startup focused on commercial applications of deep learning and computer vision. Zadeh also is the co-author of the forthcoming book TensorFlow for Deep Learning (now in early release). Our conversation took place on the eve of the recent ScaledML conference, and much of our conversation was focused on practical and real-world strategies for scaling machine learning. In particular, we spoke about the rise of deep learning, hardware/software interfaces for machine learning, and the many commercial applications of computer vision.

    57 min

About O'Reilly Data Show Podcast

From the publisher's feed

The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.

More shows like O'Reilly Data Show Podcast

Data Skeptic by Kyle Polich

Data Skeptic

476 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

624 Listeners

O'Reilly Radar Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Radar Podcast - O'Reilly Media Podcast

35 Listeners

O'Reilly Design Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Design Podcast - O'Reilly Media Podcast

8 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Machine Learning Guide by OCDevel

Machine Learning Guide

774 Listeners

DataFramed by DataCamp

DataFramed

265 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

202 Listeners

AWS Podcast by Amazon Web Services

AWS Podcast

202 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

203 Listeners

Last Week in AI by Skynet Today

Last Week in AI

316 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

99 Listeners

MIT Technology Review Narrated by MIT Technology Review

MIT Technology Review Narrated

262 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

681 Listeners

Practical News: AI & Business News by Practical News

Practical News: AI & Business News

25 Listeners