Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Keep Your Data And Query It Too Using Chaos Search with Thomas Hazel and Pete Cheslock - Episode 47
    Summary

    Elasticsearch is a powerful tool for storing and analyzing data, but when using it for logs and other time oriented information it can become problematic to keep all of your history. Chaos Search was started to make it easy for you to keep all of your data and make it usable in S3, so that you can have the best of both worlds. In this episode the CTO, Thomas Hazel, and VP of Product, Pete Cheslock, describe how they have built a platform to let you keep all of your history, save money, and reduce your operational overhead. They also explain some of the types of data that you can use with Chaos Search, how to load it into S3, and when you might want to choose it over Amazon Athena for our serverless data analysis.

    Preamble
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $/0 credit and launch a new server in under a minute.
    • You work hard to make sure that your data is reliable and accurate, but can you say the same about the deployment of your machine learning models? The Skafos platform from Metis Machine was built to give your data scientists the end-to-end support that they need throughout the machine learning lifecycle. Skafos maximizes interoperability with your existing tools and platforms, and offers real-time insights and the ability to be up and running with cloud-based production scale infrastructure instantaneously. Request a demo at dataengineeringpodcast.com/metis-machine to learn more about how Metis Machine is operationalizing data science.
    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
    • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
    • Your host is Tobias Macey and today I’m interviewing Pete Cheslock and Thomas Hazel about Chaos Search and their effort to bring historical depth to your Elasticsearch data
    • Interview
      • Introduction
      • How did you get involved in the area of data management?
      • Can you start by explaining what you have built at Chaos Search and the problems that you are trying to solve with it?
        • What types of data are you focused on supporting?
        • What are the challenges inherent to scaling an elasticsearch infrastructure to large volumes of log or metric data?

        • Is there any need for an Elasticsearch cluster in addition to Chaos Search?

        • For someone who is using Chaos Search, what mechanisms/formats would they use for loading their data into S3?

        • What are the benefits of implementing the Elasticsearch API on top of your data in S3 as opposed to using systems such as Presto or Drill to interact with the same information via SQL?

        • Given that the S3 API has become a de facto standard for many other object storage platforms, what would be involved in running Chaos Search on data stored outside of AWS?

        • What mechanisms do you use to allow for such drastic space savings of indexed data in S3 versus in an Elasticsearch cluster?

        • What is the system architecture that you have built to allow for querying terabytes of data in S3?

          • What are the biggest contributors to query latency and what have you done to mitigate them?

          • What are the options for access control when running queries against the data stored in S3?

          • What are some of the most interesting or unexpected uses of Chaos Search and access to large amounts of historical log information that you have seen?

          • What are your plans for the future of Chaos Search?

          • Contact Info
            • Pete Cheslock
              • @petecheslock on Twitter
              • Website

              • Thomas Hazel

                • @thomashazel on Twitter
                • LinkedIn

                • Parting Question
                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                  • Links
                    • Chaos Search
                    • AWS S3
                    • Cassandra
                    • Elasticsearch
                      • Podcast Interview

                      • PostgreSQL

                      • Distributed Systems

                      • Information Theory

                      • Lucene

                      • Inverted Index

                      • Kibana

                      • Logstash

                      • NVMe

                      • AWS KMS

                      • Kinesis

                      • FluentD

                      • Parquet

                      • Athena

                      • Presto

                      • Drill

                      • Backblaze

                      • OpenStack Swift

                      • Minio

                      • EMR

                      • DataDog

                      • NewRelic

                      • Elastic Beats

                      • Metricbeat

                      • Graphite

                      • Snappy

                      • Scala

                      • Akka

                      • Elastalert

                      • Tensorflow

                      • X-Pack

                      • Data Lake

                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                        Support Data Engineering Podcast

                        49 min
                      • An Agile Approach To Master Data Management with Mark Marinelli - Episode 46
                        Summary

                        With the proliferation of data sources to give a more comprehensive view of the information critical to your business it is even more important to have a canonical view of the entities that you care about. Is customer number 342 in your ERP the same as Bob Smith on Twitter? Using master data management to build a data catalog helps you answer these questions reliably and simplify the process of building your business intelligence reports. In this episode the head of product at Tamr, Mark Marinelli, discusses the challenges of building a master data set, why you should have one, and some of the techniques that modern platforms and systems provide for maintaining it.

                        Preamble
                        • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                        • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                        • You work hard to make sure that your data is reliable and accurate, but can you say the same about the deployment of your machine learning models? The Skafos platform from Metis Machine was built to give your data scientists the end-to-end support that they need throughout the machine learning lifecycle. Skafos maximizes interoperability with your existing tools and platforms, and offers real-time insights and the ability to be up and running with cloud-based production scale infrastructure instantaneously. Request a demo at dataengineeringpodcast.com/metis-machine to learn more about how Metis Machine is operationalizing data science.
                        • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                        • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                        • Your host is Tobias Macey and today I’m interviewing Mark Marinelli about data mastering for modern platforms
                        • Interview
                          • Introduction
                          • How did you get involved in the area of data management?
                          • Can you start by establishing a definition of data mastering that we can work from?
                            • How does the master data set get used within the overall analytical and processing systems of an organization?

                            • What is the traditional workflow for creating a master data set?

                              • What has changed in the current landscape of businesses and technology platforms that makes that approach impractical?
                              • What are the steps that an organization can take to evolve toward an agile approach to data mastering?

                              • At what scale of company or project does it makes sense to start building a master data set?

                              • What are the limitations of using ML/AI to merge data sets?

                              • What are the limitations of a golden master data set in practice?

                                • Are there particular formats of data or types of entities that pose a greater challenge when creating a canonical format for them?
                                • Are there specific problem domains that are more likely to benefit from a master data set?

                                • Once a golden master has been established, how are changes to that information handled in practice? (e.g. versioning of the data)

                                • What storage mechanisms are typically used for managing a master data set?

                                  • Are there particular security, auditing, or access concerns that engineers should be considering when managing their golden master that goes beyond the rest of their data infrastructure?
                                  • How do you manage latency issues when trying to reference the same entities from multiple disparate systems?

                                  • What have you found to be the most common stumbling blocks for a group that is implementing a master data platform?

                                    • What suggestions do you have to help prevent such a project from being derailed?

                                    • What resources do you recommend for someone looking to learn more about the theoretical and practical aspects of data mastering for their organization?

                                    • Contact Info
                                      • LinkedIn
                                      • Parting Question
                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                        • Links
                                          • Tamr
                                          • Multi-Dimensional Database
                                          • Master Data Management
                                          • ETL
                                          • EDW (Enterprise Data Warehouse)
                                          • Waterfall Development Method
                                          • Agile Development Method
                                          • DataOps
                                          • Feature Engineering
                                          • Tableau
                                          • Qlik
                                          • Data Catalog
                                          • PowerBI
                                          • RDBMS (Relational Database Management System)
                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                            Support Data Engineering Podcast

                                            48 min
                                          • Protecting Your Data In Use At Enveil with Ellison Anne Williams - Episode 45
                                            Summary

                                            There are myriad reasons why data should be protected, and just as many ways to enforce it in tranist or at rest. Unfortunately, there is still a weak point where attackers can gain access to your unencrypted information. In this episode Ellison Anny Williams, CEO of Enveil, describes how her company uses homomorphic encryption to ensure that your analytical queries can be executed without ever having to decrypt your data.

                                            Preamble
                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                            • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                            • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                            • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                            • Your host is Tobias Macey and today I’m interviewing Ellison Anne Williams about Enveil, a pioneering data security company protecting Data in Use
                                            • Interview
                                              • Introduction
                                              • How did you get involved in the area of data security?
                                              • Can you start by explaining what your mission is with Enveil and how the company got started?
                                              • One of the core aspects of your platform is the principal of homomorphic encryption. Can you explain what that is and how you are using it?
                                                • What are some of the challenges associated with scaling homomorphic encryption?
                                                • What are some difficulties associated with working on encrypted data sets?

                                                • Can you describe the underlying architecture for your data platform?

                                                  • How has that architecture evolved from when you first began building it?

                                                  • What are some use cases that are unlocked by having a fully encrypted data platform?

                                                  • For someone using the Enveil platform, what does their workflow look like?

                                                  • A major reason for never decrypting data is to protect it from attackers and unauthorized access. What are some of the remaining attack vectors?

                                                  • What are some aspects of the data being protected that still require additional consideration to prevent leaking information? (e.g. identifying individuals based on geographic data, or purchase patterns)

                                                  • What do you have planned for the future of Enveil?

                                                  • Contact Info
                                                    • LinkedIn
                                                    • Parting Question
                                                      • From your perspective, what is the biggest gap in the tooling or technology for data security today?
                                                      • Links
                                                        • Enveil
                                                        • NSA
                                                        • GDPR
                                                        • Intellectual Property
                                                        • Zero Trust
                                                        • Homomorphic Encryption
                                                        • Ciphertext
                                                        • Hadoop
                                                        • PII (Personally Identifiable Information)
                                                        • TLS (Transport Layer Security)
                                                        • Spark
                                                        • Elasticsearch
                                                        • Side-channel attacks
                                                        • Spectre and Meltdown
                                                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                          Support Data Engineering Podcast

                                                          25 min
                                                        • Graph Databases In Production At Scale Using DGraph with Manish Jain - Episode 44
                                                          Summary

                                                          The way that you store your data can have a huge impact on the ways that it can be practically used. For a substantial number of use cases, the optimal format for storing and querying that information is as a graph, however databases architected around that use case have historically been difficult to use at scale or for serving fast, distributed queries. In this episode Manish Jain explains how DGraph is overcoming those limitations, how the project got started, and how you can start using it today. He also discusses the various cases where a graph storage layer is beneficial, and when you would be better off using something else. In addition he talks about the challenges of building a distributed, consistent database and the tradeoffs that were made to make DGraph a reality.

                                                          Preamble
                                                          • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                          • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                          • If you have ever wished that you could use the same tools for versioning and distributing your data that you use for your software then you owe it to yourself to check out what the fine folks at Quilt Data have built. Quilt is an open source platform for building a sane workflow around your data that works for your whole team, including version history, metatdata management, and flexible hosting. Stop by their booth at JupyterCon in New York City on August 22nd through the 24th to say Hi and tell them that the Data Engineering Podcast sent you! After that, keep an eye on the AWS marketplace for a pre-packaged version of Quilt for Teams to deploy into your own environment and stop fighting with your data.
                                                          • Python has quickly become one of the most widely used languages by both data engineers and data scientists, letting everyone on your team understand each other more easily. However, it can be tough learning it when you’re just starting out. Luckily, there’s an easy way to get involved. Written by MIT lecturer Ana Bell and published by Manning Publications, Get Programming: Learn to code with Python is the perfect way to get started working with Python. Ana’s experience
                                                          • as a teacher of Python really shines through, as you get hands-on with the language without being drowned in confusing jargon or theory. Filled with practical examples and step-by-step lessons to take on, Get Programming is perfect for people who just want to get stuck in with Python. Get your copy of the book with a special 40% discount for Data Engineering Podcast listeners by going to dataengineeringpodcast.com/get-programming and use the discount code PodInit40!
                                                          • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                          • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                                          • Your host is Tobias Macey and today I’m interviewing Manish Jain about DGraph, a low latency, high throughput, native and distributed graph database.
                                                          • Interview
                                                            • Introduction
                                                            • How did you get involved in the area of data management?
                                                            • What is DGraph and what motivated you to build it?
                                                            • Graph databases and graph algorithms have been part of the computing landscape for decades. What has changed in recent years to allow for the current proliferation of graph oriented storage systems?
                                                              • The graph space is becoming crowded in recent years. How does DGraph compare to the current set of offerings?

                                                              • What are some of the common uses of graph storage systems?

                                                                • What are some potential uses that are often overlooked?

                                                                • There are a few ways that graph structures and properties can be implemented, including the ability to store data in the vertices connecting nodes and the structures that can be contained within the nodes themselves. How is information represented in DGraph and what are the tradeoffs in the approach that you chose?

                                                                • How does the query interface and data storage in DGraph differ from other options?

                                                                  • What are your opinions on the graph query languages that have been adopted by other storages systems, such as Gremlin, Cypher, and GSQL?

                                                                  • How is DGraph architected and how has that architecture evolved from when it first started?

                                                                  • How do you balance the speed and agility of schema on read with the additional application complexity that is required, as opposed to schema on write?

                                                                  • In your documentation you contend that DGraph is a viable replacement for RDBMS-oriented primary storage systems. What are the switching costs for someone looking to make that transition?

                                                                  • What are the limitations of DGraph in terms of scalability or usability?

                                                                  • Where does it fall along the axes of the CAP theorem?

                                                                  • For someone who is interested in building on top of DGraph and deploying it to production, what does their workflow and operational overhead look like?

                                                                  • What have been the most challenging aspects of building and growing the DGraph project and community?

                                                                  • What are some of the most interesting or unexpected uses of DGraph that you are aware of?

                                                                  • When is DGraph the wrong choice?

                                                                  • What are your plans for the future of DGraph?

                                                                  • Contact Info
                                                                    • @manishrjain on Twitter
                                                                    • manishrjain on GitHub
                                                                    • Blog
                                                                    • Parting Question
                                                                      • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                      • Links
                                                                        • DGraph
                                                                        • Badger
                                                                        • Google Knowledge Graph
                                                                        • Graph Theory
                                                                        • Graph Database
                                                                        • SQL
                                                                        • Relational Database
                                                                        • NoSQL
                                                                        • OLTP (On-Line Transaction Processing)
                                                                        • Neo4J
                                                                        • PostgreSQL
                                                                        • MySQL
                                                                        • BigTable
                                                                        • Recommendation System
                                                                        • Fraud Detection
                                                                        • Customer 360
                                                                        • Usenet Express
                                                                        • IPFS
                                                                        • Gremlin
                                                                        • Cypher
                                                                        • GSQL
                                                                        • GraphQL
                                                                        • MetaWeb
                                                                        • RAFT
                                                                        • Spanner
                                                                        • HBase
                                                                        • Elasticsearch
                                                                        • Kubernetes
                                                                        • TLS (Transport Layer Security)
                                                                        • Jepsen Tests
                                                                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                          Support Data Engineering Podcast

                                                                          43 min
                                                                        • Putting Airflow Into Production With James Meickle - Episode 43
                                                                          Summary

                                                                          The theory behind how a tool is supposed to work and the realities of putting it into practice are often at odds with each other. Learning the pitfalls and best practices from someone who has gained that knowledge the hard way can save you from wasted time and frustration. In this episode James Meickle discusses his recent experience building a new installation of Airflow. He points out the strengths, design flaws, and areas of improvement for the framework. He also describes the design patterns and workflows that his team has built to allow them to use Airflow as the basis of their data science platform.

                                                                          Preamble
                                                                          • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                          • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                          • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                          • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                                                          • Your host is Tobias Macey and today I’m interviewing James Meickle about his experiences building a new Airflow installation
                                                                          • Interview
                                                                            • Introduction
                                                                            • How did you get involved in the area of data management?
                                                                            • What was your initial project requirement?
                                                                              • What tooling did you consider in addition to Airflow?
                                                                              • What aspects of the Airflow platform led you to choose it as your implementation target?

                                                                              • Can you describe your current deployment architecture?

                                                                                • How many engineers are involved in writing tasks for your Airflow installation?

                                                                                • What resources were the most helpful while learning about Airflow design patterns?

                                                                                  • How have you architected your DAGs for deployment and extensibility?

                                                                                  • What kinds of tests and automation have you put in place to support the ongoing stability of your deployment?

                                                                                  • What are some of the dead-ends or other pitfalls that you encountered during the course of this project?

                                                                                  • What aspects of Airflow have you found to be lacking that you would like to see improved?

                                                                                  • What did you wish someone had told you before you started work on your Airflow installation?

                                                                                    • If you were to start over would you make the same choice?
                                                                                    • If Airflow wasn’t available what would be your second choice?

                                                                                    • What are your next steps for improvements and fixes?

                                                                                    • Contact Info
                                                                                      • @eronarn on Twitter
                                                                                      • Website
                                                                                      • eronarn on GitHub
                                                                                      • Parting Question
                                                                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                        • Links
                                                                                          • Quantopian
                                                                                          • Harvard Brain Science Initiative
                                                                                          • DevOps Days Boston
                                                                                          • Google Maps API
                                                                                          • Cron
                                                                                          • ETL (Extract, Transform, Load)
                                                                                          • Azkaban
                                                                                          • Luigi
                                                                                          • AWS Glue
                                                                                          • Airflow
                                                                                          • Pachyderm
                                                                                            • Podcast Interview

                                                                                            • AirBnB

                                                                                            • Python

                                                                                            • YAML

                                                                                            • Ansible

                                                                                            • REST (Representational State Transfer)

                                                                                            • SAML (Security Assertion Markup Language)

                                                                                            • RBAC (Role-Based Access Control)

                                                                                            • Maxime Beauchemin

                                                                                              • Medium Blog

                                                                                              • Celery

                                                                                              • Dask

                                                                                                • Podcast Interview

                                                                                                • PostgreSQL

                                                                                                  • Podcast Interview

                                                                                                  • Redis

                                                                                                  • Cloudformation

                                                                                                  • Jupyter Notebook

                                                                                                  • Qubole

                                                                                                  • Astronomer

                                                                                                    • Podcast Interview

                                                                                                    • Gunicorn

                                                                                                    • Kubernetes

                                                                                                    • Airflow Improvement Proposals

                                                                                                    • Python Enhancement Proposals (PEP)

                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                      Support Data Engineering Podcast

                                                                                                      49 min
                                                                                                    • Taking A Tour Of PostgreSQL with Jonathan Katz - Episode 42
                                                                                                      Summary

                                                                                                      One of the longest running and most popular open source database projects is PostgreSQL. Because of its extensibility and a community focus on stability it has stayed relevant as the ecosystem of development environments and data requirements have changed and evolved over its lifetime. It is difficult to capture any single facet of this database in a single conversation, let alone the entire surface area, but in this episode Jonathan Katz does an admirable job of it. He explains how Postgres started and how it has grown over the years, highlights the fundamental features that make it such a popular choice for application developers, and the ongoing efforts to add the complex features needed by the demanding workloads of today’s data layer. To cap it off he reviews some of the exciting features that the community is working on building into future releases.

                                                                                                      Preamble
                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                      • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                                                      • Are you struggling to keep up with customer request and letting errors slip into production? Want to try some of the innovative ideas in this podcast but don’t have time? DataKitchen’s DataOps software allows your team to quickly iterate and deploy pipelines of code, models, and data sets while improving quality. Unlike a patchwork of manual operations, DataKitchen makes your team shine by providing an end to end DataOps solution with minimal programming that uses the tools you love. Join the DataOps movement and sign up for the newsletter at datakitchen.io/de today. After that learn more about why you should be doing DataOps by listening to the Head Chef in the Data Kitchen at dataengineeringpodcast.com/datakitchen
                                                                                                      • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                                                      • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                                                                                      • Your host is Tobias Macey and today I’m interviewing Jonathan Katz about a high level view of PostgreSQL and the unique capabilities that it offers
                                                                                                      • Interview
                                                                                                        • Introduction
                                                                                                        • How did you get involved in the area of data management?
                                                                                                        • How did you get involved in the Postgres project?
                                                                                                        • For anyone who hasn’t used it, can you describe what PostgreSQL is?
                                                                                                          • Where did Postgres get started and how has it evolved over the intervening years?

                                                                                                          • What are some of the primary characteristics of Postgres that would lead someone to choose it for a given project?

                                                                                                            • What are some cases where Postgres is the wrong choice?

                                                                                                            • What are some of the common points of confusion for new users of PostGreSQL? (particularly if they have prior database experience)

                                                                                                            • The recent releases of Postgres have had some fairly substantial improvements and new features. How does the community manage to balance stability and reliability against the need to add new capabilities?

                                                                                                            • What are the aspects of Postgres that allow it to remain relevant in the current landscape of rapid evolution at the data layer?

                                                                                                            • Are there any plans to incorporate a distributed transaction layer into the core of the project along the lines of what has been done with Citus or CockroachDB?

                                                                                                            • What is in store for the future of Postgres?

                                                                                                            • Contact Info
                                                                                                              • @jkatz05 on Twitter
                                                                                                              • jkatz on GitHub
                                                                                                              • Parting Question
                                                                                                                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                • Links
                                                                                                                  • PostgreSQL
                                                                                                                  • Crunchy Data
                                                                                                                  • Venuebook
                                                                                                                  • Paperless Post
                                                                                                                  • LAMP Stack
                                                                                                                  • MySQL
                                                                                                                  • PHP
                                                                                                                  • SQL
                                                                                                                  • ORDBMS
                                                                                                                  • Edgar Codd
                                                                                                                  • A Relational Model of Data for Large Shared Data Banks
                                                                                                                  • Relational Algebra
                                                                                                                  • Oracle DB
                                                                                                                  • UC Berkeley
                                                                                                                  • Dr. Michael Stonebraker
                                                                                                                  • Ingres
                                                                                                                  • Informix
                                                                                                                  • QUEL
                                                                                                                  • ANSI C
                                                                                                                  • CVS
                                                                                                                  • BSD License
                                                                                                                  • UUID
                                                                                                                  • JSON
                                                                                                                  • XML
                                                                                                                  • HStore
                                                                                                                  • PostGIS
                                                                                                                  • BTree Index
                                                                                                                  • GIN Index
                                                                                                                  • GIST Index
                                                                                                                  • KNN GIST
                                                                                                                  • SPGIST
                                                                                                                  • Full Text Search
                                                                                                                  • BRIN Index
                                                                                                                  • WAL (Write-Ahead Log)
                                                                                                                  • SQLite
                                                                                                                  • PGAdmin
                                                                                                                  • Vim
                                                                                                                  • Emacs
                                                                                                                  • Linux
                                                                                                                  • OLAP (Online Analytical Processing)
                                                                                                                  • Postgres IRC
                                                                                                                  • Postgres Slack
                                                                                                                  • Postgres Conferences
                                                                                                                  • UPSERT
                                                                                                                  • Postgres Roadmap
                                                                                                                  • CockroachDB
                                                                                                                    • Podcast Interview

                                                                                                                    • Citus Data

                                                                                                                      • Podcast Interview

                                                                                                                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                        Support Data Engineering Podcast

                                                                                                                        57 min
                                                                                                                      • Mobile Data Collection And Analysis Using Ona And Canopy With Peter Lubell-Doughtie - Episode 41
                                                                                                                        Summary

                                                                                                                        With the attention being paid to the systems that power large volumes of high velocity data it is easy to forget about the value of data collection at human scales. Ona is a company that is building technologies to support mobile data collection, analysis of the aggregated information, and user-friendly presentations. In this episode CTO Peter Lubell-Doughtie describes the architecture of the platform, the types of environments and use cases where it is being employed, and the value of small data.

                                                                                                                        Preamble
                                                                                                                        • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                        • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                                                                        • Are you struggling to keep up with customer request and letting errors slip into production? Want to try some of the innovative ideas in this podcast but don’t have time? DataKitchen’s DataOps software allows your team to quickly iterate and deploy pipelines of code, models, and data sets while improving quality. Unlike a patchwork of manual operations, DataKitchen makes your team shine by providing an end to end DataOps solution with minimal programming that uses the tools you love. Join the DataOps movement and sign up for the newsletter at datakitchen.io/de today. After that learn more about why you should be doing DataOps by listening to the Head Chef in the Data Kitchen at dataengineeringpodcast.com/datakitchen
                                                                                                                        • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                                                                        • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                                                                                                        • Your host is Tobias Macey and today I’m interviewing Peter Lubell-Doughtie about using Ona for collecting data and processing it with Canopy
                                                                                                                        • Interview
                                                                                                                          • Introduction
                                                                                                                          • How did you get involved in the area of data management?
                                                                                                                          • What is Ona and how did the company get started?
                                                                                                                            • What are some examples of the types of customers that you work with?

                                                                                                                            • What types of data do you support in your collection platform?

                                                                                                                            • What are some of the mechanisms that you use to ensure the accuracy of the data that is being collected by users?

                                                                                                                            • Does your mobile collection platform allow for anyone to submit data without having to be associated with a given account or organization?

                                                                                                                            • What are some of the integration challenges that are unique to the types of data that get collected by mobile field workers?

                                                                                                                            • Can you describe the flow of the data from collection through to analysis?

                                                                                                                            • To help improve the utility of the data being collected you have started building Canopy. What was the tipping point where it became worth the time and effort to start that project?

                                                                                                                              • What are the architectural considerations that you factored in when designing it?
                                                                                                                              • What have you found to be the most challenging or unexpected aspects of building an enterprise data warehouse for general users?

                                                                                                                              • What are your plans for the future of Ona and Canopy?

                                                                                                                              • Contact Info
                                                                                                                                • Email
                                                                                                                                • pld on Github
                                                                                                                                • Website
                                                                                                                                • Parting Question
                                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                  • Links
                                                                                                                                    • OpenSRP
                                                                                                                                    • Ona
                                                                                                                                    • Canopy
                                                                                                                                    • Open Data Kit
                                                                                                                                    • Earth Institute at Columbia University
                                                                                                                                    • Sustainable Engineering Lab
                                                                                                                                    • WHO
                                                                                                                                    • Bill and Melinda Gates Foundation
                                                                                                                                    • XLSForms
                                                                                                                                    • PostGIS
                                                                                                                                    • Kafka
                                                                                                                                    • Druid
                                                                                                                                    • Superset
                                                                                                                                    • Postgres
                                                                                                                                    • Ansible
                                                                                                                                    • Docker
                                                                                                                                    • Terraform
                                                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                      Support Data Engineering Podcast

                                                                                                                                      30 min
                                                                                                                                    • Ceph: A Reliable And Scalable Distributed Filesystem with Sage Weil - Episode 40
                                                                                                                                      Summary

                                                                                                                                      When working with large volumes of data that you need to access in parallel across multiple instances you need a distributed filesystem that will scale with your workload. Even better is when that same system provides multiple paradigms for interacting with the underlying storage. Ceph is a highly available, highly scalable, and performant system that has support for object storage, block storage, and native filesystem access. In this episode Sage Weil, the creator and lead maintainer of the project, discusses how it got started, how it works, and how you can start using it on your infrastructure today. He also explains where it fits in the current landscape of distributed storage and the plans for future improvements.

                                                                                                                                      Preamble
                                                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                      • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                                                                                      • Are you struggling to keep up with customer request and letting errors slip into production? Want to try some of the innovative ideas in this podcast but don’t have time? DataKitchen’s DataOps software allows your team to quickly iterate and deploy pipelines of code, models, and data sets while improving quality. Unlike a patchwork of manual operations, DataKitchen makes your team shine by providing an end to end DataOps solution with minimal programming that uses the tools you love. Join the DataOps movement and sign up for the newsletter at datakitchen.io/de today. After that learn more about why you should be doing DataOps by listening to the Head Chef in the Data Kitchen at dataengineeringpodcast.com/datakitchen
                                                                                                                                      • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                                                                                      • Join the community in the new Zulip chat workspace at dataengineeringpodcast.com/chat
                                                                                                                                      • Your host is Tobias Macey and today I’m interviewing Sage Weil about Ceph, an open source distributed file system that supports block storage, object storage, and a file system interface.
                                                                                                                                      • Interview
                                                                                                                                        • Introduction
                                                                                                                                        • How did you get involved in the area of data management?
                                                                                                                                        • Can you start with an overview of what Ceph is?
                                                                                                                                          • What was the motivation for starting the project?
                                                                                                                                          • What are some of the most common use cases for Ceph?

                                                                                                                                          • There are a large variety of distributed file systems. How would you characterize Ceph as it compares to other options (e.g. HDFS, GlusterFS, LionFS, SeaweedFS, etc.)?

                                                                                                                                          • Given that there is no single point of failure, what mechanisms do you use to mitigate the impact of network partitions?

                                                                                                                                            • What mechanisms are available to ensure data integrity across the cluster?

                                                                                                                                            • How is Ceph implemented and how has the design evolved over time?

                                                                                                                                            • What is required to deploy and manage a Ceph cluster?

                                                                                                                                              • What are the scaling factors for a cluster?
                                                                                                                                              • What are the limitations?

                                                                                                                                              • How does Ceph handle mixed write workloads with either a high volume of small files or a smaller volume of larger files?

                                                                                                                                              • In services such as S3 the data is segregated from block storage options like EBS or EFS. Since Ceph provides all of those interfaces in one project is it possible to use each of those interfaces to the same data objects in a Ceph cluster?

                                                                                                                                              • In what situations would you advise someone against using Ceph?

                                                                                                                                              • What are some of the most interested, unexpected, or challenging aspects of working with Ceph and the community?

                                                                                                                                              • What are some of the plans that you have for the future of Ceph?

                                                                                                                                              • Contact Info
                                                                                                                                                • Email
                                                                                                                                                • @liewegas on Twitter
                                                                                                                                                • liewegas on GitHub
                                                                                                                                                • Parting Question
                                                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                  • Links
                                                                                                                                                    • Ceph
                                                                                                                                                    • Red Hat
                                                                                                                                                    • DreamHost
                                                                                                                                                    • UC Santa Cruz
                                                                                                                                                    • Los Alamos National Labs
                                                                                                                                                    • Dream Objects
                                                                                                                                                    • OpenStack
                                                                                                                                                    • Proxmox
                                                                                                                                                    • POSIX
                                                                                                                                                    • GlusterFS
                                                                                                                                                    • Hadoop
                                                                                                                                                    • Ceph Architecture
                                                                                                                                                    • Paxos
                                                                                                                                                    • relatime
                                                                                                                                                    • Prometheus
                                                                                                                                                    • Zabbix
                                                                                                                                                    • Kubernetes
                                                                                                                                                    • NVMe
                                                                                                                                                    • DNS-SD
                                                                                                                                                    • Consul
                                                                                                                                                    • EtcD
                                                                                                                                                    • DNS SRV Record
                                                                                                                                                    • Zeroconf
                                                                                                                                                    • Bluestore
                                                                                                                                                    • XFS
                                                                                                                                                    • Erasure Coding
                                                                                                                                                    • NFS
                                                                                                                                                    • Seastar
                                                                                                                                                    • Rook
                                                                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                      Support Data Engineering Podcast

                                                                                                                                                      49 min
                                                                                                                                                    • Building Data Flows In Apache NiFi With Kevin Doran and Andy LoPresto - Episode 39
                                                                                                                                                      Summary

                                                                                                                                                      Data integration and routing is a constantly evolving problem and one that is fraught with edge cases and complicated requirements. The Apache NiFi project models this problem as a collection of data flows that are created through a self-service graphical interface. This framework provides a flexible platform for building a wide variety of integrations that can be managed and scaled easily to fit your particular needs. In this episode project members Kevin Doran and Andy LoPresto discuss the ways that NiFi can be used, how to start using it in your environment, and plans for future development. They also explained how it fits in the broad landscape of data tools, the interesting and challenging aspects of the project, and how to build new extensions.

                                                                                                                                                      Preamble
                                                                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                      • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                                                                                                      • Are you struggling to keep up with customer request and letting errors slip into production? Want to try some of the innovative ideas in this podcast but don’t have time? DataKitchen’s DataOps software allows your team to quickly iterate and deploy pipelines of code, models, and data sets while improving quality. Unlike a patchwork of manual operations, DataKitchen makes your team shine by providing an end to end DataOps solution with minimal programming that uses the tools you love. Join the DataOps movement and sign up for the newsletter at datakitchen.io/de today. After that learn more about why you should be doing DataOps by listening to the Head Chef in the Data Kitchen at dataengineeringpodcast.com/datakitchen
                                                                                                                                                      • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                                                                                                      • Your host is Tobias Macey and today I’m interviewing Kevin Doran and Andy LoPresto about Apache NiFi
                                                                                                                                                      • Interview
                                                                                                                                                        • Introduction
                                                                                                                                                        • How did you get involved in the area of data management?
                                                                                                                                                        • Can you start by explaining what NiFi is?
                                                                                                                                                        • What is the motivation for building a GUI as the primary interface for the tool when the current trend is to represent everything as code?
                                                                                                                                                        • How did you get involved with the project?
                                                                                                                                                          • Where does it sit in the broader landscape of data tools?

                                                                                                                                                          • Does the data that is processed by NiFi flow through the servers that it is running on (á la Spark/Flink/Kafka), or does it orchestrate actions on other systems (á la Airflow/Oozie)?

                                                                                                                                                            • How do you manage versioning and backup of data flows, as well as promoting them between environments?

                                                                                                                                                            • One of the advertised features is tracking provenance for data flows that are managed by NiFi. How is that data collected and managed?

                                                                                                                                                              • What types of reporting are available across this information?

                                                                                                                                                              • What are some of the use cases or requirements that lend themselves well to being solved by NiFi?

                                                                                                                                                                • When is NiFi the wrong choice?

                                                                                                                                                                • What is involved in deploying and scaling a NiFi installation?

                                                                                                                                                                  • What are some of the system/network parameters that should be considered?
                                                                                                                                                                  • What are the scaling limitations?

                                                                                                                                                                  • What have you found to be some of the most interesting, unexpected, and/or challenging aspects of building and maintaining the NiFi project and community?

                                                                                                                                                                  • What do you have planned for the future of NiFi?

                                                                                                                                                                  • Contact Info
                                                                                                                                                                    • Kevin Doran
                                                                                                                                                                      • @kevdoran on Twitter
                                                                                                                                                                      • Email

                                                                                                                                                                      • Andy LoPresto

                                                                                                                                                                        • @yolopey on Twitter
                                                                                                                                                                        • Email

                                                                                                                                                                        • Parting Question
                                                                                                                                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                          • Links
                                                                                                                                                                            • NiFi
                                                                                                                                                                            • HortonWorks DataFlow
                                                                                                                                                                            • HortonWorks
                                                                                                                                                                            • Apache Software Foundation
                                                                                                                                                                            • Apple
                                                                                                                                                                            • CSV
                                                                                                                                                                            • XML
                                                                                                                                                                            • JSON
                                                                                                                                                                            • Perl
                                                                                                                                                                            • Python
                                                                                                                                                                            • Internet Scale
                                                                                                                                                                            • Asset Management
                                                                                                                                                                            • Documentum
                                                                                                                                                                            • DataFlow
                                                                                                                                                                            • NSA (National Security Agency)
                                                                                                                                                                            • 24 (TV Show)
                                                                                                                                                                            • Technology Transfer Program
                                                                                                                                                                            • Agile Software Development
                                                                                                                                                                            • Waterfall
                                                                                                                                                                            • Spark
                                                                                                                                                                            • Flink
                                                                                                                                                                            • Kafka
                                                                                                                                                                            • Oozie
                                                                                                                                                                            • Luigi
                                                                                                                                                                            • Airflow
                                                                                                                                                                            • FluentD
                                                                                                                                                                            • ETL (Extract, Transform, and Load)
                                                                                                                                                                            • ESB (Enterprise Service Bus)
                                                                                                                                                                            • MiNiFi
                                                                                                                                                                            • Java
                                                                                                                                                                            • C++
                                                                                                                                                                            • Provenance
                                                                                                                                                                            • Kubernetes
                                                                                                                                                                            • Apache Atlas
                                                                                                                                                                            • Data Governance
                                                                                                                                                                            • Kibana
                                                                                                                                                                            • K-Nearest Neighbors
                                                                                                                                                                            • DevOps
                                                                                                                                                                            • DSL (Domain Specific Language)
                                                                                                                                                                            • NiFi Registry
                                                                                                                                                                            • Artifact Repository
                                                                                                                                                                            • Nexus
                                                                                                                                                                            • NiFi CLI
                                                                                                                                                                            • Maven Archetype
                                                                                                                                                                            • IoT
                                                                                                                                                                            • Docker
                                                                                                                                                                            • Backpressure
                                                                                                                                                                            • NiFi Wiki
                                                                                                                                                                            • TLS (Transport Layer Security)
                                                                                                                                                                            • Mozilla TLS Observatory
                                                                                                                                                                            • NiFi Flow Design System
                                                                                                                                                                            • Data Lineage
                                                                                                                                                                            • GDPR (General Data Protection Regulation)
                                                                                                                                                                            • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                              Support Data Engineering Podcast

                                                                                                                                                                              1 hr 5 min
                                                                                                                                                                            • Leveraging Human Intelligence For Better AI At Alegion With Cheryl Martin - Episode 38
                                                                                                                                                                              Summary

                                                                                                                                                                              Data is often messy or incomplete, requiring human intervention to make sense of it before being usable as input to machine learning projects. This is problematic when the volume scales beyond a handful of records. In this episode Dr. Cheryl Martin, Chief Data Scientist for Alegion, discusses the importance of properly labeled information for machine learning and artificial intelligence projects, the systems that they have built to scale the process of incorporating human intelligence in the data preparation process, and the challenges inherent to such an endeavor.

                                                                                                                                                                              Preamble
                                                                                                                                                                              • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                                              • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                                                                                                                              • Are you struggling to keep up with customer request and letting errors slip into production? Want to try some of the innovative ideas in this podcast but don’t have time? DataKitchen’s DataOps software allows your team to quickly iterate and deploy pipelines of code, models, and data sets while improving quality. Unlike a patchwork of manual operations, DataKitchen makes your team shine by providing an end to end DataOps solution with minimal programming that uses the tools you love. Join the DataOps movement and sign up for the newsletter at datakitchen.io/de today. After that learn more about why you should be doing DataOps by listening to the Head Chef in the Data Kitchen at dataengineeringpodcast.com/datakitchen
                                                                                                                                                                              • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the mailing list, read the show notes, and get in touch.
                                                                                                                                                                              • Your host is Tobias Macey and today I’m interviewing Cheryl Martin, chief data scientist at Alegion, about data labelling at scale
                                                                                                                                                                              • Interview
                                                                                                                                                                                • Introduction
                                                                                                                                                                                • How did you get involved in the area of data management?
                                                                                                                                                                                • To start, can you explain the problem space that Alegion is targeting and how you operate?
                                                                                                                                                                                • When is it necessary to include human intelligence as part of the data lifecycle for ML/AI projects?
                                                                                                                                                                                • What are some of the biggest challenges associated with managing human input to data sets intended for machine usage?
                                                                                                                                                                                • For someone who is acting as human-intelligence provider as part of the workforce, what does their workflow look like?
                                                                                                                                                                                  • What tools and processes do you have in place to ensure the accuracy of their inputs?
                                                                                                                                                                                  • How do you prevent bad actors from contributing data that would compromise the trained model?

                                                                                                                                                                                  • What are the limitations of crowd-sourced data labels?

                                                                                                                                                                                    • When is it beneficial to incorporate domain experts in the process?

                                                                                                                                                                                    • When doing data collection from various sources, how do you ensure that intellectual property rights are respected?

                                                                                                                                                                                    • How do you determine the taxonomies to be used for structuring data sets that are collected, labeled or enriched for your customers?

                                                                                                                                                                                      • What kinds of metadata do you track and how is that recorded/transmitted?

                                                                                                                                                                                      • Do you think that human intelligence will be a necessary piece of ML/AI forever?

                                                                                                                                                                                      • Contact Info
                                                                                                                                                                                        • LinkedIn
                                                                                                                                                                                        • Parting Question
                                                                                                                                                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                                          • Links
                                                                                                                                                                                            • Alegion
                                                                                                                                                                                            • University of Texas at Austin
                                                                                                                                                                                            • Cognitive Science
                                                                                                                                                                                            • Labeled Data
                                                                                                                                                                                            • Mechanical Turk
                                                                                                                                                                                            • Computer Vision
                                                                                                                                                                                            • Sentiment Analysis
                                                                                                                                                                                            • Speech Recognition
                                                                                                                                                                                            • Taxonomy
                                                                                                                                                                                            • Feature Engineering
                                                                                                                                                                                            • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                                              Support Data Engineering Podcast

                                                                                                                                                                                              47 min

                                                                                                                                                                                            About Data Engineering Podcast

                                                                                                                                                                                            From the publisher's feed

                                                                                                                                                                                            This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

                                                                                                                                                                                            More shows like Data Engineering Podcast

                                                                                                                                                                                            This Week in Startups by Jason Calacanis

                                                                                                                                                                                            This Week in Startups

                                                                                                                                                                                            1,289 Listeners

                                                                                                                                                                                            The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                                                                            The Changelog: Software Development, Open Source

                                                                                                                                                                                            286 Listeners

                                                                                                                                                                                            The a16z Show by Andreessen Horowitz

                                                                                                                                                                                            The a16z Show

                                                                                                                                                                                            1,089 Listeners

                                                                                                                                                                                            Software Engineering Daily by Software Engineering Daily

                                                                                                                                                                                            Software Engineering Daily

                                                                                                                                                                                            622 Listeners

                                                                                                                                                                                            Risky Business by Risky Business Media

                                                                                                                                                                                            Risky Business

                                                                                                                                                                                            374 Listeners

                                                                                                                                                                                            Talk Python To Me by Michael Kennedy

                                                                                                                                                                                            Talk Python To Me

                                                                                                                                                                                            582 Listeners

                                                                                                                                                                                            Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                                                                            Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                                                                            304 Listeners

                                                                                                                                                                                            NVIDIA AI Podcast by NVIDIA

                                                                                                                                                                                            NVIDIA AI Podcast

                                                                                                                                                                                            337 Listeners

                                                                                                                                                                                            Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                                                                            Syntax - Tasty Web Development Treats

                                                                                                                                                                                            985 Listeners

                                                                                                                                                                                            Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                                                                            Practical AI

                                                                                                                                                                                            203 Listeners

                                                                                                                                                                                            Dwarkesh Podcast by Dwarkesh Patel

                                                                                                                                                                                            Dwarkesh Podcast

                                                                                                                                                                                            565 Listeners

                                                                                                                                                                                            The Data Engineering Show by The Firebolt Data Bros

                                                                                                                                                                                            The Data Engineering Show

                                                                                                                                                                                            8 Listeners

                                                                                                                                                                                            Latent Space: The AI Engineer Podcast by Latent.Space

                                                                                                                                                                                            Latent Space: The AI Engineer Podcast

                                                                                                                                                                                            102 Listeners

                                                                                                                                                                                            This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                                                                                                                                                                                            This Day in AI Podcast

                                                                                                                                                                                            222 Listeners

                                                                                                                                                                                            The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                                                                                                                                                                                            The AI Daily Brief: Artificial Intelligence News and Analysis

                                                                                                                                                                                            685 Listeners