O'Reilly Data Show Podcast

O'Reilly Data Show Podcast

By O'Reilly Media
Download on the App Store

O'Reilly Data Show Podcast episodes

  • Bringing scalable real-time analytics to the enterprise

    In this episode of the Data Show, I spoke with Dhruba Borthakur (co-founder and CTO) and Shruti Bhat (SVP of Product) of Rockset, a startup focused on building solutions for interactive data science and live applications. Borthakur was the founding engineer of HDFS and creator of RocksDB, while Bhat is an experienced product and marketing executive focused on enterprise software and data products. Their new startup is focused on a few trends I’ve recently been thinking about, including the re-emergence of real-time analytics, and the hunger for simpler data architectures and tools.  Borthakur exemplifies the need for companies to continually evaluate new technologies: while he was the founding engineer for HDFS, these days he mostly works with object stores like S3.

    38 min
  • Applications of data science and machine learning in financial services

    In this episode of the Data Show, I spoke with Jike Chong, chief data scientist at Acorns, a startup focused on building tools for micro-investing. Chong has extensive experience using analytics and machine learning in financial services, and he has experience building data science teams in the U.S. and in China.

    We had a great conversation spanning many topics, including:

    • Potential applications of data science in financial services.
    • The current state of data science in financial services in both the U.S. and China.
    • His experience recruiting, training, and managing data science teams in both the U.S. and China.
    • Here are some highlights from our conversation:

      Opportunities in financial services

      There’s a customer acquisition piece and then there’s a customer retention piece. For customer acquisition, we can see that new technologies can really add value by looking at all sorts of data sources that can help a financial service company identify who they want to target to provide those services. So, it’s a great place where data science can help find the product market fit, not just at one instance like identifying who you want to target, but also in a continuous form where you can evolve a product and then continuously find the audience that would best fit the product and continue to analyze the audience so you can design the next generation product. … Once you have a specific cohort of users who you want to target, there’s a need to be able to precisely convert them, which means understanding the stage of the customer’s thought process and understanding how to form the narrative to convince the user or the customer that a particular piece of technology or particular piece of service is the current service they need.

      … On the customer serving or retention side, for financial services we commonly talk about building hundred-year businesses, right? They have to be profitable businesses, and for financial service to be profitable, there are operational considerations—quantifying risk requires a lot of data science; preventing fraud is really important, and there is garnering the long-term trust with the customer so they stay with you, which means having the work ethic to be able to take care of customer’s data and able to serve the customer better with automated services whenever and wherever the customer is. It’s all those opportunities where I see we can help serve the customer by having the right services presented to them and being able to serve them in the long term.

      Opportunities in China

      A few important areas in the financial space in China include mobile payments, wealth management, lending, and insurance—basically, the major areas for the financial industry.

      For these areas, China may be a forerunner in using internet technologies, especially mobile internet technologies for FinTech, and I think the wave started way back in the 2012/2013 time frame. If you look at mobile payments, like Alipay and WeChat, those have hundreds of millions of active users. The latest data from Alipay is about 608 million users, and these are monthly active users we’re talking about. This is about two times the U.S. population actively using Alipay on a monthly basis, which is a crazy number if you consider all the data that can generate and all the things you can see people buying to be able to understand how to serve the users better.

      If you look at WeChat, they’re boasting one billion users, monthly active users, early this year. Those are the huge players, and with that amount of traffic, they are able to generate a lot of interest for the lower-frequency services like wealth management and lending, as well as insurance.

      Related resources:

      • Kai-Fu Lee outlines the factors that enabled China’s rapid ascension in AI
      • Gary Kazantsev on how “Data science makes an impact on Wall Street”
      • Juan Huerta on “Upcoming challenges and opportunities for data technologies in consumer finance”
      • Geoffrey Bradway on “Programming collective intelligence for financial trading”
      • Jason Dai on why “Companies in China are moving quickly to embrace AI technologies”
      • Haoyuan Li on why “In the age of AI, fundamental value resides in data”
      • 43 min
      • Real-time entity resolution made accessible

        In this episode of the Data Show, I spoke with Jeff Jonas, CEO, founder and chief scientist of Senzing, a startup focused on making real-time entity resolution technologies broadly accessible. He was previously a fellow and chief scientist of context computing at IBM. Entity resolution (ER) refers to techniques and tools for identifying and linking manifestations of the same entity/object/individual. Ironically, ER itself has many different names (e.g., record linkage, duplicate detection, object consolidation/reconciliation, etc.).

        ER is an essential first step in many domains, including marketing (cleaning up databases), law enforcement (background checks and counterterrorism), and financial services and investing. Knowing exactly who your customers are is an important task for security, fraud detection, marketing, and personalization. The proliferation of data sources and services has made ER very challenging in the internet age. In addition, many applications now increasingly require near real-time entity resolution.

        We had a great conversation spanning many topics including:

        • Why ER is interesting and challenging
        • How ER technologies have evolved over the years
        • How Senzing is working to democratize ER by making real-time AI technologies accessible to developers
        • Some early use cases for Senzing’s technologies
        • Some items on their research agenda
        • Here are a few highlights from our conversation:

          Entity Resolution through years

          In the early ’90s, I worked on a much more advanced version of entity resolution for the casinos in Las Vegas and created software called NORA, non-obvious relationship awareness. Its purpose was to help casinos better understand who they were doing business with. We would ingest data from the loyalty club, everybody making hotel reservations, people showing up without reservations, everybody applying for jobs, people terminated, vendors, and 18 different lists of different kinds of bad people, some of them card counters (which aren’t that bad), some cheaters. And they wanted to figure out across all these identities when somebody was the same, and then when people were related. Some people were using 32 different names and a bunch of different social security numbers.

          … Ultimately, IBM bought my company and this technology became what is known now at IBM as “identity insight.” Identity insight is a real-time entity resolution engine that gets used to solve many kinds of problems. MoneyGram implemented it and their fraud complaints dropped 72%. They saved a few hundred million just in their first few years.

          … But while at IBM, I had a grand vision about a new type of entity resolution engine that would have been unlike anything that’s ever existed. It’s almost like a Swiss Army knife for ER.

          Recent developments

          The Senzing entity resolution engine works really well on two records from a domain that you’ve never even seen before. Say you’ve never done entity resolution on restaurants from Singapore. The first two records you feed it, it’s really, really already smart. And then as you feed it more data, it gets smarter and smarter.

          … So, there are two things that we’ve intertwined. One is common sense. One type of common sense is the names—Dick, Dickie, Richie, Rick, Ricardo are all part of the same name family. Why should it have to study millions and millions of records to learn that again?

          … Next to common sense, there’s real-time learning. In real-time learning, we do a few things. You might have somebody named Bob, but who now goes by a nickname or an alias of Andy. Eventually, you might come to learn that. So, now you know you have to learn over time that Bob also has this nickname, and Bob lived at three addresses, and this is his credit card number, and now he’s got four phone numbers. So you want to learn those over time.

          … These systems we’re creating, our entity resolution systems—which really resolve entities and graph them (call it index of identities and how they’re related)—never has to be reloaded. It literally cleans itself up in the past. You can do maintenance on it while you’re querying it, while you’re loading new transactional data, while you’re loading historical data. There’s nothing else like it that can work at this scale. It’s really hard to do.

          Related resources:

          • Jeff Jonas on “Context Computing”
          • David Ferrucci on why “Language understanding remains one of AI’s grand challenges”
          • David Blei on “Topic models: Past, present, and future”
          • “Lessons learned building natural language processing systems in health care”
          • “Building a contacts graph from activity data”
          • “Customer record deduplication using Spark and Reifier”
          • 28 min
          • Why companies are in need of data lineage solutions

            In this episode of the Data Show, I spoke with Neelesh Salian, software engineer at Stitch Fix, a company that combines machine learning and human expertise to personalize shopping. As companies integrate machine learning into their products and systems, there are important foundational technologies that come into play. This shouldn’t come as a shock, as current machine learning and AI technologies require large amounts of data—specifically, labeled data for training models. There are also many other considerations—including security, privacy, reliability/safety—that are encouraging companies to invest in a suite of data technologies. In conversations with data engineers, data scientists, and AI researchers, the need for solutions that can help track data lineage and provenance keeps popping up.

            There are several San Francisco Bay Area companies that have embarked on building data lineage systems—including Salian and his colleagues at Stitch Fix. I wanted to find out how they arrived at the decision to build such a system and what capabilities they are building into it.

            Here are some highlights from our conversation:

            Data lineage

            Data lineage is not something new. It’s something that is borne out of the necessity of understanding how data is being written and interacted with in the data warehouse. I like to tell this story when I’m describing data lineage: think of it as a journey for data. The data takes a journey entering into your warehouse. This can be transactional data, dashboards, or recommendations. What is lost in that collection of data is the information about how it came about. If you knew what journey and exactly what constituted that data to come into being into your data warehouse or any other storage appliance you use, that would be really useful.

            … Think about data lineage as helping issues about quality of data, understanding if something is corrupted. On the security side, think of GDPR … which was one of the hot topics I heard about at the Strata Data Conference in London in 2018.

            Why companies are suddenly building data lineage solutions

            A data lineage system becomes necessary as time progresses. It becomes easier for maintainability. You need it for audit trails, for security and compliance. But you also need to think of the benefit of managing the data sets you’re working with. If you’re working with 10 databases, you need to know what’s going on in them. If I have to give you a vision of a data lineage system, think of it as a final graph or view of some data set, and it shows you a graph of what it’s linked to. Then it gives you some metadata information so you can drill down. Let’s say you have corrupted data, let’s say you want to debug something. All these cases tie into the actual use cases for which we want to build it.

            Related resources:

            • “Deep automation in machine learning”
            • Vitaly Gordon on “Building tools for enterprise data science”
            • “Managing risk in machine learning”
            • Haoyuan Li explains why “In the age of AI, fundamental value resides in data”
            • “What machine learning means for software development”
            • Joe Hellerstein on how “Metadata services can lead to performance and organizational improvements”
            • 35 min
            • What data scientists and data engineers can do with current generation serverless technologies

              In this episode of the Data Show, I spoke with Avner Braverman, co-founder and CEO of Binaris, a startup that aims to bring serverless to web-scale and enterprise applications. This conversation took place shortly after the release of a seminal paper from UC Berkeley (“Cloud Programming Simplified: A Berkeley View on Serverless Computing”), and this paper seeded a lot of our conversation during this episode.

              Serverless is clearly on the radar of data engineers and architects. In a recent survey, we found 85% of respondents already had parts of their data infrastructure in one of the public clouds, and 38% were already using at least one of the serverless offerings we listed. As more serverless offerings get rolled out—e.g., things like PyWren that target scientists—I expect these numbers to rise.

              We had a great conversation spanning many topics, including:

              • A short history of cloud computing.
              • The fundamental differences between serverless and conventional cloud computing.
              • The reasons serverless—specifically AWS Lambda—took off so quickly.
              • What can data scientists and data engineers do with the current generation serverless offerings.
              • What is missing from serverless today and what should users expect in the near future.
              • Related resources:

                • “The evolution and expanding utility of Ray”
                • Results of a new survey: “Evolving Data Infrastructure: Tools and Best Practices for Advanced Analytics and AI”
                • Eric Jonas on “Building accessible tools for large-scale computation and machine learning”
                • “7 data trends on our radar”
                • “Handling real-time data operations in the enterprise”
                • “Progress for big data in Kubernetes”
                • 37 min
                • It’s time for data scientists to collaborate with researchers in other disciplines

                  In this episode of the Data Show, I spoke with Forough Poursabzi-Sangdeh, a postdoctoral researcher at Microsoft Research New York City. Poursabzi works in the interdisciplinary area of interpretable and interactive machine learning. As models and algorithms become more widespread, many important considerations are becoming active research areas: fairness and bias, safety and reliability, security and privacy, and Poursabzi’s area of focus—explainability and interpretability.

                  We had a great conversation spanning many topics, including:

                  • Current best practices and state-of-the-art methods used to explain or interpret deep learning—or, more generally, machine learning models.
                  • The limitations of current model interpretability methods.
                  • The lack of clear/standard metrics for comparing different approaches used for model interpretability
                  • Many current AI and machine learning applications augment humans, and, thus, Poursabzi believes it’s important for data scientists to work closely with researchers in other disciplines.
                  • The importance of using human subjects in model interpretability studies.
                  • Related resources:

                    • “Local Interpretable Model-Agnostic Explanations (LIME): An Introduction”
                    • “Interpreting predictive models with Skater: Unboxing model opacity”
                    • Jacob Ward on “How social science research can inform the design of AI systems”
                    • Sharad Goel and Sam Corbett-Davies on “Why it’s hard to design fair machine learning models”
                    • “Managing risk in machine learning”: considerations for a world where ML models are becoming mission critical
                    • Francesca Lazzeri and Jaya Mathew on “Lessons learned while helping enterprises adopt machine learning”
                    • Jerry Overton on “Teaching and implementing data science and AI in the enterprise”
                    • 37 min
                    • Algorithms are shaping our lives—here’s how we wrest back control
                      In this episode of the Data Show, I spoke with Kartik Hosanagar, professor of technology and digital business, and professor of marketing at The Wharton School of the University of Pennsylvania.  Hosanagar is also the author of a newly released book, A Human’s Guide to Machine Intelligence, an interesting tour through the recent evolution of AI applications that draws from his extensive experience at the intersection of business and technology.
                      We had a great conversation spanning many topics, including:
                      The types of unanticipated consequences of which algorithm designers should be aware.
                      The predictability-resilience paradox: as systems become more intelligent and dynamic, they also become more unpredictable, so there are trade-offs algorithms designers must face.
                      Managing risk in machine learning: AI application designers need to weigh considerations such as fairness, security, privacy, explainability, safety, and reliability.
                      A bill of rights for humans impacted by the growing power and sophistication of algorithms.
                      Some best practices for bringing AI into the enterprise.
                      Related resources:
                      “Managing risk in machine learning”: considerations for a world where ML models are becoming mission critical
                      Francesca Lazzeri and Jaya Mathew on “Lessons learned while helping enterprises adopt machine learning”
                      Jerry Overton on “Teaching and implementing data science and AI in the enterprise”
                      Kris Hammond on “Bringing AI into the enterprise”
                      Jacob Ward on “How social science research can inform the design of AI systems”
                      “Overcoming barriers to AI adoption”
                      Sharad Goel and Sam Corbett-Davies on “Why it’s hard to design fair machine learning models”
                      45 min
                    • Why your attention is like a piece of contested territory
                      In this episode of the Data Show, I spoke with P.W. Singer, strategist and senior fellow at the New America Foundation, and a contributing editor at Popular Science. He is co-author of an excellent new book, LikeWar: The Weaponization of Social Media, which explores how social media has changed war, politics, and business. The book is essential reading for anyone interested in how social media has become an important new battlefield in a diverse set of domains and settings.
                      We had a great conversation spanning many topics, including:
                      In light of the 10th anniversary of his earlier book Wired for War, we talked about progress in robotics over the past decade.
                      The challenge posed by the fact that social networks reward virality, not veracity.
                      How the internet has emerged as an important new battlefield.
                      How this new online battlefield changes how conflicts are fought and unfold.
                      How many of the ideas and techniques covered in LikeWar are trickling down from nation-state actors influencing global events, to consulting companies offering services that companies and individuals can use.
                      Here are some highlights from our conversation:
                      LikeWar
                      We spent five years tracking how social media was being used all around the world. … We looked at everything from how was it being used by militaries, by terrorist groups, by politicians, by teenagers—you name it. The finding of this project is sort of a two-fold play on words. The first is, if you think of cyberwar as the hacking of networks, LikeWar is its twin. It’s the hacking of people on the networks by driving ideas viral through a mix of likes and lies.
                      … Social media began as a space for fun, for entertainment. It then became a communication space. It became a marketplace. It’s also turned it into a kind of battle space. It’s simultaneously all of these things at once, and you can see, for example, Russian information warriors who are using digital marketing techniques and teenage jokes to influence the outcomes of elections. A different example would be ISIS’ top recruiter, Junaid Hussain, mimicking how Taylor Swift built her fan army.
                      A common set of tactics
                      The second finding of the project was that when you look across all these wildly diverse actors, groups, and organizations, they turned out to be using very similar tactics, very similar approaches. To put it a different way: it’s a mode of conflict. There’s ways of “winning” that all the different groups are realizing. More importantly, the groups that understand these new rules of the game are the ones that are winning their online wars and having a real effect, whether that real effect is winning a political campaign, winning a corporate marketing campaign, winning a campaign to become a celebrity, or to become the most popular kid in school. Or “winning” might be to do the opposite—to sabotage someone else’s campaign to become a leading political candidate.
                      Related resources:
                      Siwei Lyu on “The technical, societal, and cultural challenges that come with the rise of fake media”
                      Supasorn Suwajanakorn on “Building artificial people: Endless possibilities and the dark side”
                      Guillaume Chaslot on “The importance of transparency and user control in machine learning”
                      “Overcoming barriers to AI adoption”
                      Alon Kaufman on “Machine learning on encrypted data”
                      Sharad Goel and Sam Corbett-Davies on “Why it’s hard to design fair machine learning models”
                      44 min
                    • The technical, societal, and cultural challenges that come with the rise of fake media
                      In this episode of the Data Show, I spoke with Siwei Lyu, associate professor of computer science at the University at Albany, State University of New York. Lyu is a leading expert in digital media forensics, a field of research into tools and techniques for analyzing the authenticity of media files. Over the past year, there have been many stories written about the rise of tools for creating fake media (mainly images, video, audio files). Researchers in digital image forensics haven’t exactly been standing still, though. As Lyu notes, advances in machine learning and deep learning have also found a receptive audience among the forensics community.
                      We had a great conversation spanning many topics including:
                      The many indicators used by forensic experts and forgery detection systems
                      Balancing “open” research with risks that come with it—including “tipping off” adversaries
                      State-of-the-art detection tools today, and what the research community and funding agencies are working on over the next few years.
                      Technical, societal, and cultural challenges that come with the rise of fake media.
                      Here are some highlights from our conversation:
                      Imbalance between digital forensics researchers and forgers
                      In theory, it looks difficult to synthesize media. This is true, but on the other hand, there are factors to consider on the side of the forgers. The first is the fact that most people working in forensics, like myself, usually just write a paper and publish it. So, the details of our detection algorithm becomes available immediately. On the other hand, people making fake media are usually secretive; they don’t usually publish the details of their algorithms. So, there’s a kind of imbalance between the information on the forensic side and the forgery side.
                      The other issue is user habit. The fact that even if some of the fakes are very low quality, a typical user checks it just for a second; sees something interesting, exciting, sensational; and helps distribute it without actually checking the authenticity. This actually helps fake media to broadcast very, very fast. Even though we have algorithms to detect fake media, these tools are probably not fast enough to actually stop the trap.
                      … Then there are the actual incentives for this kind of work. For forensics, even if we have the tools and the time to catch a piece of fake media, we don’t get anything. But for people actually making the fake media, there is more financial or other forms of incentive to do that.
                      Related resources:
                      Supasorn Suwajanakorn on “Building artificial people: Endless possibilities and the dark side”
                      Alyosha Efros on “Using computer vision to understand big visual data”
                      “Overcoming barriers to AI adoption”
                      “What is neural architecture search?”
                      Alon Kaufman on “Machine learning on encrypted data”
                      Sharad Goel and Sam Corbett-Davies on “Why it’s hard to design fair machine learning models”
                      31 min
                    • Using machine learning and analytics to attract and retain employees
                      In this episode of the Data Show, I spoke with Maryam Jahanshahi, research scientist at TapRecruit, a startup that uses machine learning and analytics to help companies recruit more effectively. In an upcoming survey, we found that a “skills gap” or “lack of skilled people” was one of the main bottlenecks holding back adoption of AI technologies. Many companies are exploring a variety of internal and external programs to train staff on new tools and processes. The other route is to hire new talent. But recent reports suggest that demand for data professionals is strong and competition for experienced talent is fierce. Jahanshahi and her team are building natural language and statistical tools that can help companies improve their ability to attract and retain talent across many key areas.
                      Here are some highlights from our conversation:
                      Optimal job titles
                      The conventional wisdom in our field has always been that you want to optimize for “the number of good candidates” divided by “the number of total candidates.” … The thinking is that one of the ways in which you get a good signal-to-noise ratio is if you advertise for a more senior role. … In fact, we found the number of qualified applicants was lower for the senior data scientist role.
                      … We saw from some of our behavioral experiments that people were feeling like that was too senior a role for them to apply to. What we would call the “confidence gap” was kicking in at that point. It’s a pretty well-known phenomena that there are different groups of the population that are less confident. This has been best characterized in terms of gender. It’s the idea that most women only apply for jobs when they meet 100% of the qualifications versus most men will apply even with 60% of the qualifications. That was actually manifesting.
                      Highlighting benefits
                      We saw a lot of big companies that would offer 401(k), that would offer health insurance or family leave, but wouldn’t mention those benefits in the job descriptions. This had an impact on how candidates perceived these companies. Even though it’s implied that Coca-Cola is probably going to give you 401(k) and health insurance, not mentioning it changes the way you think of that job.
                      … So, don’t forget the things that really should be there. Even the boring stuff really matters for most candidates. You’d think it would only matter for older candidates, but, actually, millennials and everyone in every age group are very concerned about these things because it’s not specifically about the 401(k) plan; it’s about what it implies in terms of the company—that the company is going to take care of you, is going to give you leave, is going to provide a good workplace.
                      Improving diversity
                      We found the best way to deal with representation at the end of the process is actually to deal with representation early in the process. What I mean by that is having a robust or a healthy candidate pool at the start of the process. We found for data scientist roles, that was about having 100 candidates apply for your job.
                      … If we’re not getting to the point where we can attract 100 applicants, we’ll take a look at that job description. We’ll see what’s wrong with it and what could be turning off candidates; it could be that you’re not syndicating the job description well, it’s not getting into search results, or it could be that it’s actually turning off a lot of people. You could be asking for too many qualifications, and that turns off a lot of people. … Sometimes it involves taking a step back and taking a look at what we’re doing in this process that’s not helping us and that’s starving us of candidates.
                      Related resources:
                      Sharad Goel and Sam Corbett-Davies on “Why it’s hard to design fair machine learning models”
                      “Comparing production-grade NLP libraries”
                      “What are machine learning engineers?”
                      “Th
                      47 min

                    About O'Reilly Data Show Podcast

                    From the publisher's feed

                    The O'Reilly Data Show Podcast explores the opportunities and techniques driving big data, data science, and AI.

                    More shows like O'Reilly Data Show Podcast

                    Data Skeptic by Kyle Polich

                    Data Skeptic

                    476 Listeners

                    Software Engineering Daily by Software Engineering Daily

                    Software Engineering Daily

                    624 Listeners

                    O'Reilly Radar Podcast - O'Reilly Media Podcast by O'Reilly Media

                    O'Reilly Radar Podcast - O'Reilly Media Podcast

                    35 Listeners

                    O'Reilly Design Podcast - O'Reilly Media Podcast by O'Reilly Media

                    O'Reilly Design Podcast - O'Reilly Media Podcast

                    8 Listeners

                    Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                    Super Data Science: ML & AI Podcast with Jon Krohn

                    305 Listeners

                    NVIDIA AI Podcast by NVIDIA

                    NVIDIA AI Podcast

                    338 Listeners

                    Machine Learning Guide by OCDevel

                    Machine Learning Guide

                    774 Listeners

                    DataFramed by DataCamp

                    DataFramed

                    265 Listeners

                    Practical AI by Daniel Whitenack and Chris Benson

                    Practical AI

                    202 Listeners

                    AWS Podcast by Amazon Web Services

                    AWS Podcast

                    202 Listeners

                    Google DeepMind: The Podcast by Hannah Fry

                    Google DeepMind: The Podcast

                    203 Listeners

                    Last Week in AI by Skynet Today

                    Last Week in AI

                    316 Listeners

                    Machine Learning Street Talk (MLST) by Machine Learning Street Talk (MLST)

                    Machine Learning Street Talk (MLST)

                    99 Listeners

                    MIT Technology Review Narrated by MIT Technology Review

                    MIT Technology Review Narrated

                    263 Listeners

                    This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                    This Day in AI Podcast

                    222 Listeners

                    The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                    The AI Daily Brief: Artificial Intelligence News and Analysis

                    681 Listeners

                    Practical News: AI & Business News by Practical News

                    Practical News: AI & Business News

                    25 Listeners