O'Reilly Data Show Podcast

O'Reilly Data Show Podcast

By O'Reilly Media
Download on the App Store

O'Reilly Data Show Podcast episodes

  • The importance of transparency and user control in machine learning
    In this episode of the Data Show, I spoke with Guillaume Chaslot, an ex-YouTube engineer and founder of AlgoTransparency, an organization dedicated to helping the public understand the profound impact algorithms have on our lives. We live in an age when many of our interactions with companies and services are governed by algorithms. At a time when their impact continues to grow, there are many settings where these algorithms are far from transparent. There is growing awareness about the vast amounts of data companies are collecting on their users and customers, and people are starting to demand control over their data. A similar conversation is starting to happen about algorithms—users are wanting more control over what these models optimize for and an understanding of how they work.
    I first came across Chaslot through a series of articles about the power and impact of YouTube on politics and society. Many of the articles I read relied on data and analysis supplied by Chaslot. We talked about his work trying to decipher how YouTube’s recommendation system works, filter bubbles, transparency in machine learning, and data privacy.
    Here are some highlights from our conversation:
    Why YouTube’s impact is less understood
    My theory why people completely overlooked YouTube is because on Facebook and Twitter, if one of your friends posts something strange, you’ll see it. Even if you have 1,000 friends, if one of them posts something really disturbing, you see it, so you’re more aware of the problem. Whereas on YouTube, some people binge watch some very weird things that could be propaganda, but we won’t know about it because we don’t see what other people see. So, YouTube is like a TV channel that doesn’t show the same thing to everybody and when you ask YouTube, “What did you show to other people?” YouTube says, ‘I don’t know, I don’t remember, I don’t want to tell you.’
    Downsides of optimizing only for watch time
    When I was working on the YouTube algorithm and our goal was to optimize watch time, we were trying to make sure that the algorithm kept people online the longest. But what I realized was that we were so focused on this target of watch time that we were forgetting a lot of important things and we were seeing some very strange behavior of the algorithm. Each time we were seeing this strange behavior, we just blamed it on the user. It shows violent videos; it must be because users are violent, so it’s not our fault; the algorithm is just a mirror of human society. But if I believe the algorithm is a mirror of human society, I think it’s also not a flat mirror; it’s a mirror that emphasizes some aspects of life and makes some other aspects overlooked.
    … The algorithm that is behind YouTube and the Facebook news feeds are very complex, deep learning systems that will take a lot into account, including user sessions, what they’ve watched. It will try to find the right content to show to users to get them to stay online the longest and interact as much as possible with the content. So, this can seem neutral at first, but it might not be neutral. For instance, if you have content that says ‘The media is lying,’ whether it’s on Facebook or on YouTube, what will happen is that this content will naturally, if it manages to convince the user that the media is lying, the content will be very efficient at keeping the user online because the user won’t go to other media and will spend more time on YouTube and more time on Facebook.
    … In my personal opinion, the current goal of maximizing watch time means that any content that is really good at captivating your attention for a long time will perform really well. This means extreme content will actually perform really well. But say you had another goal—for instance, the goal to maximize likes and dislikes, or another system of rating like when you would be asked some question like, ‘Did yo
    24 min
  • What machine learning engineers need to know
    In this episode of the Data Show, I spoke with Jesse Anderson, managing director of the Big Data Institute, and my colleague Paco Nathan, who recently became co-chair of Jupytercon. This conversation grew out of a recent email thread the three of us had on machine learning engineers, a new job role that LinkedIn recently pegged as the fastest growing job in the U.S. In our email discussion, there was some disagreement on whether such a specialized job role/title was needed in the first place. As Eric Colson pointed out in his beautiful keynote at Strata Data San Jose, when done too soon, creating specialized roles can slow down your data team.
    We recorded this conversation at Strata San Jose, while Anderson was in the middle of teaching his very popular two-day training course on real-time systems. We closed the conversation with Anderson’s take on Apache Pulsar, a very impressive new messaging system that is starting to gain fans among data engineers.
    Here are some highlights from our conversation:
    Why we need machine learning engineers
    Jesse Anderson: (2:09) One of the issues I’m seeing as I work with teams is that they’re trying to operationalize machine learning models, and the data scientists are not the one to productionize these. They simply don’t have the engineering skills. Conversely, the data engineers don’t have the skills to operationalize this either. So, we’re seeing this kind of gap in between the data science and the data engineering, and the gap I’m seeing and the way I’m seeing it being filled, is through a machine learning engineer.
    … I disagree with Paco that generalization is the way to go. I think it’s hyper-specialization, actually. This is coming from my experience having taught a lot of enterprises. At a startup, I would say that super-specialization is probably not going to be as possible, but at an enterprise, you are going to have to have a team that specializes in big data, and that is a part from a team, even a software engineering team, that doesn’t work with data.
    Putting Apache Pulsar on the radar of data engineers
    Key features of Apache Pulsar. Image by Karthik Ramasamy, used with permission.
    Jesse Anderson: (23:30) A lot of my time, since I’m really teaching data engineering is spent on data integration and data ingestion. How do we move this data around efficiently? For a lot of that time Kafka was really the only open source game in town for that. But now there’s another technology called Apache Pulsar. I’ve spent a decent amount of time actually going through Pulsar and there are some things that I see in it that Kafka will either have difficulty doing or won’t be able to do.
    … Apache Pulsar separates pub-sub from storage. When I first read about that, I didn’t quite get it. I didn’t quite see, why is this so important or why this is so interesting. It’s because you can individually scale your pub-sub and your storage resources independently. Now you’ve got something. Now you can say, “Well, we originally decided I wanted to store data for seven days. All right, let’s spin up some more bookkeeper processes and now we can store fourteen days, now we can store twenty one days.” I think that’s going to be a pretty interesting addition there. Where the other side of that, the corollary to that is, “Okay, we’re hitting Black Friday and we don’t have so much more data coming through as we have way more consumption and have way more things hitting our pub-sub. We could spin up more pub-sub with that.” This separation is actually allowing some interesting use cases.
    Related resources:
    “What are machine learning engineers?”
    “We need to build machine learning tools to augment machine learning engineers”
    “Differentiating via data science”: Eric Colson explains why companies must now think very differently abou
    33 min
  • How to train and deploy deep learning at scale
    In this episode of the Data Show, I spoke with Ameet Talwalkar, assistant professor of machine learning at CMU and co-founder of Determined AI. He was an early and key contributor to Spark MLlib and a member of AMPLab. Most recently, he helped conceive and organize the first edition of SysML, a new academic conference at the intersection of systems and machine learning (ML).
    We discussed using and deploying deep learning at scale. This is an empirical era for machine learning, and, as I noted in an earlier article, as successful as deep learning has been, our level of understanding of why it works so well is still lacking. In practice, machine learning engineers need to explore and experiment using different architectures and hyperparameters before they settle on a model that works for their specific use case. Training a single model usually involves big (labeled) data and big models; as such, exploring the space of possible model architectures and parameters can take days, weeks, or even months. Talwalkar has spent the last few years grappling with this problem as an academic researcher and as an entrepreneur. In this episode, he describes some of his related work on hyperparameter tuning, systems, and more.
    Here are some highlights from our conversation:
    Deep learning
    I would say that you hear a lot about the modeling of problems associated with deep learning. How do I frame my problem as a machine learning problem? How do I pick my architecture? How do I debug things when things go wrong? … What we’ve seen in practice is that, maybe somewhat surprisingly, the biggest challenges that ML engineers face actually are due to the lack of tools and software for deep learning. These problems are sort of like hybrid systems/ML problems. Very similar to the sorts of research that came out of the AMPLab.
    … Things like TensorFlow and Keras, and a lot of those other platforms that you mentioned, are great and they’re a great step forward. They’re really good at abstracting low-level details of a particular learning architecture. In five lines, you can describe how your architecture looks and then you can also specify what algorithms you want to use for training.
    There are a lot of other systems challenges associated with actually going end to end, from data to a deployed model. The existing software solutions don’t really tackle a big set of these challenges. For example, regardless of the software you’re using, it takes days to weeks to train a deep learning model. There’s real open challenges of how to best use parallel and distributed computing both to train a particular model and in the context of tuning hyperparameters of different models.
    We also found out the vast majority of organizations that we’ve spoken to in the last year or so who are using deep learning for what I’d call mission-critical problems, are actually doing it with on-premise hardware. Managing this hardware is a huge challenge and something that folks like me, if I’m working at a company with machine learning engineers, have to figure out for themselves. It’s kind of a mismatch between their interests and their skills, but it’s something they have to take care of.
    Understanding distributed training
    To give a little bit more background, the idea behind this work started about four years ago. There was no deep learning in Spark MLlib at the time. We were trying to figure out how to perform distributed training of deep learning in Spark. Before actually getting our hands really dirty and trying to actually implement anything we wanted to just do some back-of-the-envelope calculations to see what speed-ups you could hope to get.
    … The two main ingredients here are just computation and communication. … We wanted to understand this landscape of distributed training, and, using Paleo, we’ve been able to get a good sense of this landscape without actually running experiments. The i
    40 min
  • Using machine learning to monitor and optimize chatbots
    In this episode of the Data Show, I spoke with Ofer Ronen, GM of Chatbase, a startup housed within Google’s Area 120. With tools for building chatbots becoming accessible, conversational interfaces are becoming more prevalent. As Ronen highlights in our conversation, chatbots are already enabling companies to automate many routine tasks (mainly in customer interaction). We are still in the early days of chatbots, but if current trends persist, we’ll see bots deployed more widely and take on more complex tasks and interactions. Gartner recently predicted that by 2021, companies will spend more on bots and chatbots than mobile app development.
    Like any other software application, as bots get deployed in real-world applications, companies will need tools to monitor their performance. For a single, simple chatbot, one can imagine developers manually monitoring log files for errors and problems. Things get harder as you scale to more bots and as the bots get increasingly more complex. As in the case of other machine learning applications, when companies start deploying many more chatbots, automated tools for monitoring and diagnostics become essential.
    The good news is relevant tools are beginning to emerge. In this episode, Ronen describes a tool he helped build: Chatbase is a chatbot analytics and optimization service that leverages machine learning research and technologies developed at Google. In essence, Chatbase lets companies focus on building and deploying the best possible chatbots.
    Here are some highlights from our conversation:
    Democratization of tools for bot developers
    It’s been hard to get the natural language processing to work well and to recognize all the different ways people might say the same thing. There’s been an explosion of tools that leverage machine learning and natural language processing (NLP) engines to make sense of all that’s being asked of bots. But with increased capacity and capability to process data, there’s now better third-party tools for any company to take advantage of and build a decent bot out of the box.
    … I see three levels of bot builders out there. There’s the non-technical kind where marketing or sales might create a prototype using a user interface—like maybe Chatfuel, which requires no programming, and create a basic experience. Or they might even create some sort of decision tree bot that is not flexible, but is good for maybe basic lead-gen experiences. But they often can’t handle type-ins. It’s often button-based. So, that’s one level, the non-technical folks.
    Then there are teams that have developers on staff. They’re not machine learning experts, but they’re developers that can use off-the-shelf natural language processing engines to extract meaning from messages sent by users. So, you’re extracting intents and entities and making sense of what’s coming at your bot without having to have machine learning expertise.
    Finally, there are teams that have the machine learning experts. They might build their own NLP engine to give them more control over how it works. But often, that’s not needed if a third-party solution can serve most of your needs. But we do see some teams like that.
    Popular use cases
    We track tens and tens of thousands of bots each month with Chatbase, and what we see is that for large companies, they often start with customer support for at least two reasons. One is that the automation can save them some money, but also because chatbots enable them to create a more effective experience for their users. In fact, there’s a survey by Salesforce and a couple other companies that found that what people want from bots is quick 24/7 answers to simple questions.
    We also see some lead-generation bots. Those are simpler to build and often just live on a website of a company to try to gather and qualify leads. They can, most of the time, do a decent job and do a little better than just
    28 min
  • Unleashing the potential of reinforcement learning
    In this episode of the Data Show, I spoke with Danny Lange, VP of AI and machine learning at Unity Technologies. Lange previously led data and machine learning teams at Microsoft, Amazon, and Uber, where his teams were responsible for building data science tools used by other developers and analysts within those companies. When I first heard that he was moving to Unity, I was curious as to why he decided to join a company whose core product targets game developers.
    As you’ll glean from our conversation, Unity is at the forefront of some of the most exciting, practical applications of deep learning (DL) and reinforcement learning (RL). Realistic scenery and imagery are critical for modern games. GANs and related semi-supervised techniques can ease content creation by enabling artists to produce realistic images much more quickly. In a previous post, Lange described how reinforcement learning opens up the possibility of training/learning rather than programming in game development.
    Lange explains why simulation environments are going to be important tools for AI developers. We are still in the early days of machine intelligence, and I am looking forward to more tools that can democratize AI research (including future releases by Lange and his team at Unity).
    Here are some highlights from our conversation:
    Why reinforcement learning is so exciting
    I’m a huge fan of reinforcement learning. I think it has incredible potential, not just in game development but in a lot of other areas, too. … What we are doing at Unity is basically making reinforcement learning available to the masses. We have shipped open source software on GitHub called Unity ML Agents, that include the basic frameworks for people to experiment with reinforcement learning. Reinforcement learning is really creating a machine learned-driven feedback loop. Recall the example I previously wrote about, of the chicken crossing the road; yes, it gets hit thousands and thousands of times by these cars, but every time it gets hit, it learns that’s a bad thing. And every time it manages to pick up a gift package on the way over the road, that’s a good thing.
    Over time, it gets superhuman capabilities in crossing this road, and that is fantastic because there’s not a single line of code going into that. It’s pure simulation, and through reinforcement learning it captures a method. It learns a method to cross the road, and you can take that into many different aspects of games. There are many different methods you can train. You can add two chickens—can they collaborate to do something together? We are looking at what we call multi-agent systems, where two or more of these trained reinforcement learning-trained agents are acting together to achieve a goal.
    … I want a million developers to start working on this. I want a lot more innovation, and I want a lot more out-of-the-box thinking, and that is what we want by making our RL tools and platform available to our Unity community. Let me just jump to one thing here: most people think that reinforcement learning in the game world or in game-like situations is a lot about what we call ‘path finding.’ Path finding is basically for a character in a game to navigate through some situation—this is pretty well understood. There are good algorithms for that. Looking ahead, I’m actually thinking about a different set of decisions. For instance, which weapon or which tool should a character pick up and bring with them in a game? That is a much, much harder decision. It’s strategy at a higher level.
    Machine learning and AI at Unity
    If you think about where intelligence originated around us (animals and humans), it’s really originating out of surviving and thriving in a physical world. That is really the job of intelligence. You have to survive, you have to find food, you have to avoid your enemies, you have to walk falling down—so, gravity is playing a big role there. If you thi
    34 min
  • Graphs as the front end for machine learning
    In this episode of the Data Show, I spoke with Leo Meyerovich, co-founder and CEO of Graphistry. Graphs have always been part of the big data revolution (think of the large graphs generated by the early social media startups). In recent months, I’ve come across companies releasing and using new tools for creating, storing, and (most importantly) analyzing large graphs. There are many problems and use cases that lend themselves naturally to graphs, and recent advances in hardware and software building blocks have made large-scale analytics possible.
    Starting with his work as a graduate student at UC Berkeley, Meyerovich has pioneered the combination of hardware and software acceleration to create truly interactive environments for visualizing large amounts of data. Graphistry has built a suite of tools that enables analysts to wade through large data sets and investigate business and security incidents. The company is currently focused on the security domain—where it turns out that graph representations of data are things security analysts are quite familiar with.
    Here are some highlights from our conversation:
    Graphs as the front end for machine learning
    They’re really flexible. First of all, there’s a pure analytic reason in that there are certain types of queries that one could do efficiently with a graph database. If you needed do a bunch of joins, graphs are really great at that. … Companies want to get into stuff like 360-degree views of things; they want to understand correlations to actually explain what’s going on at a more intelligent level.
    … I think that’s where graphs really start to shine. Because companies deal with pretty heterogeneous data, and a graph ends up being a really easy way to deal with that. A lot of questions are basically, “What’s nearby?”—almost like your nearest neighbor type of stuff; the graph becomes, both at the query level and at the visual level, very interpretable. I now have a hypothesis about graphs as being the front end and the UI for machine learning, but that might be a topic for another day.
    Graph applications and correlation services
    If we’re talking about investigating a financial crime, you’ve got a transaction or user. … For example, if the user has multiple names but all the names are using the same address, you’re going to want to see that relationship.
    … In security, where a lot of my mind is today, there is something called Kill Chain, where if you think of any bad incident, there’s probably a sequence of events around it, that led up to it. … You can map out that Kill Chain. So, in a sense, a lot of the reason Graphistry uses graphs is so we can let people see that sort of progression of events and reason about it.
    … When people are using the graphs, especially in an enterprise setting, I think there’s a process change that’s happening if you’re building an enterprise data lake type of system. … It’s great if you can get individual alerts and create cases and investigations around individual alerts. But increasingly, you want a higher-level thing. … Instead of looking at individual alerts or individual events, you really want to think of incidents—an incident is basically a collection of alerts. For example, maybe there’s some fraud going on; if somebody figured out how to do fraud once, they’re probably going to try doing it multiple times. So, you don’t want to be playing whack-a-mole on little symptoms in each individual case; you want to get that full group incident. A graph becomes basically a way to create a real correlation service.
    (Full disclosure: I’m an advisor to Graphistry.)
    Related resources:
    “Graph databases are powering mission-critical applications”: Emil Eifrem on popular applications of graph technologies
    “Semi-supervised, unsupervised, and adaptive algorithms for large-scale time series”: Ira Cohen on developi
    46 min
  • Machine learning needs machine teaching
    In this episode of the Data Show, I spoke with Mark Hammond, founder and CEO of Bonsai, a startup at the forefront of developing AI systems in industrial settings. While many articles have been written about developments in computer vision, speech recognition, and autonomous vehicles, I’m particularly excited about near-term applications of AI to manufacturing, robotics, and industrial automation. In a recent post, I outlined practical applications of reinforcement learning (RL)—a type of machine learning now being used in AI systems. In particular, I described how companies like Bonsai are applying RL to manufacturing and industrial automation. As researchers explore new approaches for solving RL problems, I expect many of the first applications to be in industrial automation.
    Here are some highlights from our conversation:
    Machine learning and machine teaching
    Everyone is so focused on making better and faster learning algorithms; what do we do when we have it? Let’s just suppose that you now have an algorithm that can learn as well as or better than humans. How do we use that, how do we apply that in a predictable, scalable, repeatable way toward the objectives that we want to apply it toward?
    … I thought about that for a while, and it’s one of those things where the answer is obvious in hindsight, but until you sit down and really chew on it, it doesn’t jump out at you. And it’s that, by design, if you’re building a learning system—if you want to program it—you have to teach it. Machine teaching and machine learning are necessary complements to one another; you need both. And for the large part, most of what comprises machine teaching these days consists of giant label data sets.
    … You need machine teaching and machine learning. It dawned on me that this was the core abstraction that was going to make it possible for us to start applying all of this stuff more broadly across all the myriad use cases that we see in the real world without having to turn all of the people who are looking to use it into experts in machine learning and data science. It’s what enabled me to realize what Bonsai’s mission is: to enable your subject matter experts (a chemical engineer or a mechanical engineer, someone who is very, very well versed in whatever their domain is but not necessarily in machine learning or data science) to take that expertise and use it as the foundation for describing what to teach and then automating the underlying pieces for how you can actually effectively learn that.
    Common use cases for reinforcement learning
    In a lot of cases, the system you are working with is undergoing some form of tuning. We see a lot of this. It might be that you have an HVAC system and you are looking to optimize across the lifetime of the equipment, the comfort of the user, and the energy consumption of the device, which is its own complex optimization problem. You do this by virtue of tuning a bunch of parameters or controlling a bunch of different capabilities on the device.
    … So, in many cases, a lot of what our users are applying the system toward is not necessarily strict real-time control. Certainly, we do that as well, and we put out a paper recently showing dexterous robotic control and manipulation using these kinds of techniques. You can absolutely do that, and we have customers who do that. A very common thing that we see across a lot of the real-world applications is more a tuning characteristic.
    Related resources:
    “Practical applications of reinforcement learning in industry”: an overview of commercial and industrial applications of reinforcement learning
    “Introducing RLlib – A composable and scalable reinforcement learning library”: this new software makes the task of training RL models much more accessible
    Deep reinforcement learning in the enterprise—Bridging the gap from games to industry (2017 Artificial Intelligence Conference presentation by Mark Hammond)
    Ray
    46 min
  • How machine learning can be used to write more secure computer programs
    In this episode of the Data Show, I spoke with Fabian Yamaguchi, chief scientist at ShiftLeft. His 2015 Ph.D. dissertation sketched out how the combination of static analysis, graph mining, and machine learning, can be used to develop tools to augment security analysts. In a recent post, I argued for machine learning tools to augment teams responsible for deploying and managing models in production (machine learning engineers). These are part of a general trend of using machine learning to develop and manage the software systems of tomorrow. Yamaguchi’s work is step one in this direction: using machine learning to reduce the number of security vulnerabilities in complex software products.
    Here are some highlights from our conversation:
    Machine learning to find code vulnerabilities
    I was not trying to build something that would just automatically take the code and give you all of the vulnerabilities. Instead, I was looking at the typical kind of tasks that I would encounter myself when doing these security audits, and I would ask myself, how can I automate these subtasks? As an example, when you find a vulnerability in code, the question that often arises is whether there are similar vulnerabilities still in that same program. That’s one of those subtasks you can automate well because what you’re actually doing is saying: ‘Hey, here’s an example of what a bug looks like. Can you scan the rest of the code? Can you use machine learning to actually determine other locations in the code that implement the same bug?’
    … In machine learning, you never have enough data. In this case, this is actually an unsupervised learning approach. You’re taking all of the functions that you can get and you extract the dominant programming patterns in there. … It’s a bit like what you would do to find similar text documents, but it’s used for code.
    From source code to graph analytics
    By transforming software code into a graph, you can actually extract different properties from that code by analyzing the graph.
    … Let’s take a smaller function that might have one IF block. One of the graph structures that’s first generated is called an abstract syntax tree. That’s a tree that you’d get by just parsing the code. …  For each IF and for each variable, for each statement, there’s going to be a node. For each operator, like if there’s an assignment, there’s also going to be a node, and they are all connected by edges. You soon run into a lot of nodes and edges. If you take something like, let’s say, the Linux kernel, you’ll have several hundreds of thousands of nodes.
    … You can do a lot by essentially solving reachability problems in these graphs.
    Related resources:
    “How machine learning will accelerate data management systems”
    Artificial intelligence in the software engineering workflow: A 2017 AI Conference keynote by Peter Norvig
    “Responsible deployment of machine learning”: Why we need to build machine learning tools to augment our machine learning engineers
    “Architecting and building end-to-end streaming applications”
    Data is only as valuable as the decisions it enables
    29 min
  • Bringing AI into the enterprise
    In this episode of the Data Show, I spoke with Kristian Hammond, chief scientist of Narrative Science and professor of EECS at Northwestern University. He has been at the forefront of helping companies understand the power, limitations, and disruptive potential of AI technologies and tools. In a previous post on machine learning, I listed types of uses cases (a taxonomy) for machine learning that could just as well apply to enterprise applications of AI. But how do you identify good use cases to begin with?
    A good place to start for most companies is by looking for AI technologies that can help automate routine tasks, particularly low-skill tasks that occupy the time of high-skilled workers. An initial list of candidate tasks can be gathered by applying the following series of simple questions:
    Is the task data-driven?
    Do you have the data to support the automation of the task?
    Do you really need the scale that automation can provide?
    We discussed other factors companies should consider when thinking through their AI strategies, education and training programs for AI specialists, and the importance of ethics and fairness in AI and data science.
    Here are some highlights from our conversation:
    It begins with finding use cases
    I’ve been interacting more and more with the companies that are thinking about AI solutions; they often won’t have gotten to the place where they can talk about what they want to do. It’s an odd thing because there’s so much data out there and there’s so much hunger to derive something from that data. The starting point is often bringing an organization back down to, “So what do you want and need to do? What kind of decision-making do you want to support? What kinds of predictions would you like to be able to make?”
    Identifying which tasks can be automated
    Sometimes, you see a decision being made and, from an organizational point of view, everyone agrees that this decision is really strongly data driven. But it’s not strongly data driven. It’s data driven based upon the historical information that two or three people are using. It looks like they’re looking at data and then making a decision, but, in fact, what they’re doing is, they’re looking at data and they’re remembering one of 2,000 past examples in their heads and coming out with a decision.
    … There are sets of tasks in almost any organization that nobody likes to have anything to do with. In the legal profession, there are tasks around things like discovery where you actually need to be able to look through a corpus of documents, but you need to have also some idea of the semantic relationships between words. This is totally learnable using existing technologies.
    … It’s not as though tasks that can be automated don’t exist. They do, and, in fact, they not only exist, but they’re easily doable with current technologies. It’s a matter of understanding where to draw the line. It’s sometimes easy for organizations to look at the problem and sort of hallucinate that there is not a different kind of reasoning going on in the heads of the people who are solving the problem.
    … You have to be willing to look at that and say, “Oh, I’m not going to replace the smartest person in the company, but, you know, I will free up the time of some of our smartest people by taking these tasks on and having the machine do them.”
    Related resources:
    Here and now – Bringing AI into the enterprise: Kris Hammond’s tutorial at the 2017 AI conference in San Francisco.
    Vertical AI – Solving full stack industry problems using subject-matter expertise, unique data, and AI to deliver a product’s core value proposition: Bradford Cross at the 2017 AI conference in San Francisco.
    Demystifying the AI hype: Kathryn Hume at the 2017 AI conference in NYC.
    “6 practical guidelines for implementing conversational AI“: Susan
    45 min
  • How machine learning will accelerate data management systems
    In this episode of the Data Show, I spoke with Tim Kraska, associate professor of computer science at MIT. To take advantage of big data, we need scalable, fast, and efficient data management systems. Database administrators and users often find themselves tasked with building index structures (“indexes” in database parlance), which are needed to speed up data access.
    Some common examples include:
    B-Trees—used for range requests (e.g., assemble all sales orders within a certain time frame)
    Hash maps—used for key-based lookups
    Bloom filters—used to check whether an element or piece of data is present in a set
    Index structures take up space in a database, so you need to be selective about what to index, and they do not take advantage of the underlying data distributions. I’ve worked in settings where an administrator or expert user carefully implements a strategy for building indexes for a data warehouse based on important and common queries.
    Indexes are really models or mappings—for instance, a Bloom filter can be thought of as a classification problem. In a recent paper, Kraska and his collaborators approach indexing as a learning problem. As such, they are able to build indexes that take into account underlying data distributions, are smaller in size (thus allowing for a more liberal indexing strategy), and their indexes execute faster. Software and hardware for computation are getting cheaper and better, so using machine learning to create index structures is something that may indeed become routine.
    This ties with a larger trend of using machine learning to improve software systems and even software development. In the future, we’ll have database administrators who have machine learning tools at their disposal, which would allow them to manage larger and more complex systems, and these ML tools will free them to focus on complex tasks that are harder to automate.
    Here are some highlights from our conversation:
    Why use machine learning to learn index structures
    I think it used to be the case that if you know you have the key distribution, you could leverage that, but you need to build a very specialized system for that. Then, if the data distribution changes, you need to adjust the whole system. At the same time, any learning mechanism was in the past and way too expensive to do it.
    Things have changed a little bit because compute is becoming much cheaper. Suddenly, using machine learning to train this mapping actually pays off. On one hand, the B-tree structures are composed of a whole bunch of “IF statements,” and in the past, multiplications were very expensive. Now multiplications are getting cheaper and cheaper. Scaling “IF statements” is hard, but scaling math operations is at least relatively easier. In essence, we can trade these “IF statements” for multiplications, and that’s actually why suddenly learning the data distribution pays off.
    … For B-trees, for example, we saw speed-ups of up to roughly 2X. However, the indexes were up to two orders of magnitude smaller.
    Why B-Trees are models. Image from Tim Kraska, used with permission.
    The future of data management systems
    If this machine learning approach really works out, I think this might change the way database systems are built. … Maybe the database administrator (DBA) of the future becomes a machine learning expert.
    … My hope is that the system can figure out what model to use, but maybe if you want to have the best performance and you know your data very well, I can see that maybe the DBA / machine learning expert chooses a certain type of model to tune the index.
    … There was this Tweet, essentially saying that machine learning will change how we build core algorithms and data structures. I think this is currently still the better analogy.
    Related resources:
    Tupleware—redefining modern data analytics: a Strata Data 2014 presentation by Tim Kraska
    Artificial intelligence in the software engineering workflow: a 2017 AI Conference keyn
    35 min

About O'Reilly Data Show Podcast

From the publisher's feed

The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.

More shows like O'Reilly Data Show Podcast

Data Skeptic by Kyle Polich

Data Skeptic

476 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

624 Listeners

O'Reilly Radar Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Radar Podcast - O'Reilly Media Podcast

35 Listeners

O'Reilly Design Podcast - O'Reilly Media Podcast by O'Reilly Media

O'Reilly Design Podcast - O'Reilly Media Podcast

8 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Machine Learning Guide by OCDevel

Machine Learning Guide

774 Listeners

DataFramed by DataCamp

DataFramed

265 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

202 Listeners

AWS Podcast by Amazon Web Services

AWS Podcast

202 Listeners

Google DeepMind: The Podcast by Hannah Fry

Google DeepMind: The Podcast

203 Listeners

Last Week in AI by Skynet Today

Last Week in AI

316 Listeners

Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

Machine Learning Street Talk (MLST)

99 Listeners

MIT Technology Review Narrated by MIT Technology Review

MIT Technology Review Narrated

263 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

681 Listeners

Practical News: AI & Business News by Practical News

Practical News: AI & Business News

25 Listeners