Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Pulsar: Fast And Scalable Messaging with Rajan Dhabalia and Matteo Merli - Episode 17
    Summary

    One of the critical components for modern data infrastructure is a scalable and reliable messaging system. Publish-subscribe systems have been popular for many years, and recently stream oriented systems such as Kafka have been rising in prominence. This week Rajan Dhabalia and Matteo Merli discuss the work they have done on Pulsar, which supports both options, in addition to being globally scalable and fast. They explain how Pulsar is architected, how to scale it, and how it fits into your existing infrastructure.

    Preamble
    • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
    • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
    • You can help support the show by checking out the Patreon page which is linked from the site.
    • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
    • A few announcements:
      • There is still time to register for the O’Reilly Strata Conference in San Jose, CA March 5th-8th. Use the link dataengineeringpodcast.com/strata-san-jose to register and save 20%
      • The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
      • If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.

      • Your host is Tobias Macey and today I’m interviewing Rajan Dhabalia and Matteo Merli about Pulsar, a distributed open source pub-sub messaging system

      • Interview
        • Introduction
        • How did you get involved in the area of data management?
        • Can you start by explaining what Pulsar is and what the original inspiration for the project was?
        • What have been some of the most challenging aspects of building and promoting Pulsar?
        • For someone who wants to run Pulsar, what are the infrastructure and network requirements that they should be considering and what is involved in deploying the various components?
        • What are the scaling factors for Pulsar and what aspects of deployment and administration should users pay special attention to?
        • What projects or services do you consider to be competitors to Pulsar and what makes it stand out in comparison?
        • The documentation mentions that there is an API layer that provides drop-in compatibility with Kafka. Does that extend to also supporting some of the plugins that have developed on top of Kafka?
        • One of the popular aspects of Kafka is the persistence of the message log, so I’m curious how Pulsar manages long-term storage and reprocessing of messages that have already been acknowledged?
        • When is Pulsar the wrong tool to use?
        • What are some of the improvements or new features that you have planned for the future of Pulsar?
        • Contact Info
          • Matteo
            • merlimat on GitHub
            • @merlimat on Twitter

            • Rajan

              • @dhabaliaraj on Twitter
              • rhabalia on GitHub

              • Parting Question
                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                • Links
                  • Pulsar
                  • Publish-Subscribe
                  • Yahoo
                  • Streamlio
                  • ActiveMQ
                  • Kafka
                  • Bookkeeper
                  • SLA (Service Level Agreement)
                  • Write-Ahead Log
                  • Ansible
                  • Zookeeper
                  • Pulsar Deployment Instructions
                  • RabbitMQ
                  • Confluent Schema Registry
                    • Podcast Interview

                    • Kafka Connect

                    • Wallaroo

                      • Podcast Interview

                      • Kinesis

                      • Athenz

                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                        Support Data Engineering Podcast

                        54 min
                      • Dat: Distributed Versioned Data Sharing with Danielle Robinson and Joe Hand - Episode 16
                        Summary

                        Sharing data across multiple computers, particularly when it is large and changing, is a difficult problem to solve. In order to provide a simpler way to distribute and version data sets among collaborators the Dat Project was created. In this episode Danielle Robinson and Joe Hand explain how the project got started, how it functions, and some of the many ways that it can be used. They also explain the plans that the team has for upcoming features and uses that you can watch out for in future releases.

                        Preamble
                        • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                        • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                        • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                        • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                        • You can help support the show by checking out the Patreon page which is linked from the site.
                        • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                        • A few announcements:
                          • There is still time to register for the O’Reilly Strata Conference in San Jose, CA March 5th-8th. Use the link dataengineeringpodcast.com/strata-san-jose to register and save 20%
                          • The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
                          • If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.
                          • Your host is Tobias Macey and today I’m interviewing Danielle Robinson and Joe Hand about Dat Project, a distributed data sharing protocol for building applications of the future
                          • Interview
                            • Introduction
                            • How did you get involved in the area of data management?
                            • What is the Dat project and how did it get started?
                            • How have the grants to the Dat project influenced the focus and pace of development that was possible?
                              • Now that you have established a non-profit organization around Dat, what are your plans to support future sustainability and growth of the project?
                              • Can you explain how the Dat protocol is designed and how it has evolved since it was first started?
                              • How does Dat manage conflict resolution and data versioning when replicating between multiple machines?
                              • One of the primary use cases that is mentioned in the documentation and website for Dat is that of hosting and distributing open data sets, with a focus on researchers. How does Dat help with that effort and what improvements does it offer over other existing solutions?
                              • One of the difficult aspects of building a peer-to-peer protocol is that of establishing a critical mass of users to add value to the network. How have you approached that effort and how much progress do you feel that you have made?
                              • How does the peer-to-peer nature of the platform affect the architectural patterns for people wanting to build applications that are delivered via dat, vs the common three-tier architecture oriented around persistent databases?
                              • What mechanisms are available for content discovery, given the fact that Dat URLs are private and unguessable by default?
                              • For someone who wants to start using Dat today, what is involved in creating and/or consuming content that is available on the network?
                              • What have been the most challenging aspects of building and promoting Dat?
                              • What are some of the most interesting or inspiring uses of the Dat protocol that you are aware of?
                              • Contact Info
                                • Dat
                                  • datproject.org
                                  • Email
                                  • @dat_project on Twitter
                                  • Dat Chat
                                  • Danielle
                                    • Email
                                    • @daniellecrobins
                                    • Joe
                                      • Email
                                      • @joeahand on Twitter
                                      • Parting Question
                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                        • Links
                                          • Dat Project
                                          • Code For Science and Society
                                          • Neuroscience
                                          • Cell Biology
                                          • OpenCon
                                          • Mozilla Science
                                          • Open Education
                                          • Open Access
                                          • Open Data
                                          • Fortune 500
                                          • Data Warehouse
                                          • Knight Foundation
                                          • Alfred P. Sloan Foundation
                                          • Gordon and Betty Moore Foundation
                                          • Dat In The Lab
                                          • Dat in the Lab blog posts
                                          • California Digital Library
                                          • IPFS
                                          • Dat on Open Collective – COMING SOON!
                                          • ScienceFair
                                          • Stencila
                                          • eLIFE
                                          • Git
                                          • BitTorrent
                                          • Dat Whitepaper
                                          • Merkle Tree
                                          • Certificate Transparency
                                          • Dat Protocol Working Group
                                          • Dat Multiwriter Development – Hyperdb
                                          • Beaker Browser
                                          • WebRTC
                                          • IndexedDB
                                          • Rust
                                          • C
                                          • Keybase
                                          • PGP
                                          • Wire
                                          • Zenodo
                                          • Dryad Data Sharing
                                          • Dataverse
                                          • RSync
                                          • FTP
                                          • Globus
                                          • Fritter
                                          • Fritter Demo
                                          • Rotonde how to
                                          • Joe’s website on Dat
                                          • Dat Tutorial
                                          • Data Rescue – NYTimes Coverage
                                          • Data.gov
                                          • Libraries+ Network
                                          • UC Conservation Genomics Consortium
                                          • Fair Data principles
                                          • hypervision
                                          • hypervision in browser
                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                            Click here to read the unedited transcript…
                                            Tobias Macey 00:13…
                                            1 hr 3 min
                                          • Snorkel: Extracting Value From Dark Data with Alex Ratner - Episode 15
                                            Summary

                                            The majority of the conversation around machine learning and big data pertains to well-structured and cleaned data sets. Unfortunately, that is just a small percentage of the information that is available, so the rest of the sources of knowledge in a company are housed in so-called “Dark Data” sets. In this episode Alex Ratner explains how the work that he and his fellow researchers are doing on Snorkel can be used to extract value by leveraging labeling functions written by domain experts to generate training sets for machine learning models. He also explains how this approach can be used to democratize machine learning by making it feasible for organizations with smaller data sets than those required by most tooling.

                                            Preamble
                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                            • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                            • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                            • You can help support the show by checking out the Patreon page which is linked from the site.
                                            • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                            • Your host is Tobias Macey and today I’m interviewing Alex Ratner about Snorkel and Dark Data
                                            • Interview
                                              • Introduction
                                              • How did you get involved in the area of data management?
                                              • Can you start by sharing your definition of dark data and how Snorkel helps to extract value from it?
                                              • What are some of the most challenging aspects of building labelling functions and what tools or techniques are available to verify their validity and effectiveness in producing accurate outcomes?
                                              • Can you provide some examples of how Snorkel can be used to build useful models in production contexts for companies or problem domains where data collection is difficult to do at large scale?
                                              • For someone who wants to use Snorkel, what are the steps involved in processing the source data and what tooling or systems are necessary to analyse the outputs for generating usable insights?
                                              • How is Snorkel architected and how has the design evolved over its lifetime?
                                              • What are some situations where Snorkel would be poorly suited for use?
                                              • What are some of the most interesting applications of Snorkel that you are aware of?
                                              • What are some of the other projects that you and your group are working on that interact with Snorkel?
                                              • What are some of the features or improvements that you have planned for future releases of Snorkel?
                                              • Contact Info
                                                • Website
                                                • ajratner on Github
                                                • @ajratner on Twitter
                                                • Parting Question
                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                  • Links
                                                    • Stanford
                                                    • DAWN
                                                    • HazyResearch
                                                    • Snorkel
                                                    • Christopher Ré
                                                    • Dark Data
                                                    • DARPA
                                                    • Memex
                                                    • Training Data
                                                    • FDA
                                                    • ImageNet
                                                    • National Library of Medicine
                                                    • Empirical Studies of Conflict
                                                    • Data Augmentation
                                                    • PyTorch
                                                    • Tensorflow
                                                    • Generative Model
                                                    • Discriminative Model
                                                    • Weak Supervision
                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                      Support Data Engineering Podcast

                                                      38 min
                                                    • CRDTs and Distributed Consensus with Christopher Meiklejohn - Episode 14
                                                      Summary

                                                      As we scale our systems to handle larger volumes of data, geographically distributed users, and varied data sources the requirement to distribute the computational resources for managing that information becomes more pronounced. In order to ensure that all of the distributed nodes in our systems agree with each other we need to build mechanisms to properly handle replication of data and conflict resolution. In this episode Christopher Meiklejohn discusses the research he is doing with Conflict-Free Replicated Data Types (CRDTs) and how they fit in with existing methods for sharing and sharding data. He also shares resources for systems that leverage CRDTs, how you can incorporate them into your systems, and when they might not be the right solution. It is a fascinating and informative treatment of a topic that is becoming increasingly relevant in a data driven world.

                                                      Preamble
                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                      • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                      • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                      • You can help support the show by checking out the Patreon page which is linked from the site.
                                                      • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                      • Your host is Tobias Macey and today I’m interviewing Christopher Meiklejohn about establishing consensus in distributed systems
                                                      • Interview
                                                        • Introduction
                                                        • How did you get involved in the area of data management?
                                                        • You have dealt with CRDTs with your work in industry, as well as in your research. Can you start by explaining what a CRDT is, how you first began working with them, and some of their current manifestations?
                                                        • Other than CRDTs, what are some of the methods for establishing consensus across nodes in a system and how does increased scale affect their relative effectiveness?
                                                        • One of the projects that you have been involved in which relies on CRDTs is LASP. Can you describe what LASP is and what your role in the project has been?
                                                        • Can you provide examples of some production systems or available tools that are leveraging CRDTs?
                                                        • If someone wants to take advantage of CRDTs in their applications or data processing, what are the available off-the-shelf options, and what would be involved in implementing custom data types?
                                                        • What areas of research are you most excited about right now?
                                                        • Given that you are currently working on your PhD, do you have any thoughts on the projects or industries that you would like to be involved in once your degree is completed?
                                                        • Contact Info
                                                          • Website
                                                          • cmeiklejohn on GitHub
                                                          • Google Scholar Citations
                                                          • Parting Question
                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                            • Links
                                                              • Basho
                                                              • Riak
                                                              • Syncfree
                                                              • LASP
                                                              • CRDT
                                                              • Mesosphere
                                                              • CAP Theorem
                                                              • Cassandra
                                                              • DynamoDB
                                                              • Bayou System (Xerox PARC)
                                                              • Multivalue Register
                                                              • Paxos
                                                              • RAFT
                                                              • Byzantine Fault Tolerance
                                                              • Two Phase Commit
                                                              • Spanner
                                                              • ReactiveX
                                                              • Tensorflow
                                                              • Erlang
                                                              • Docker
                                                              • Kubernetes
                                                              • Erleans
                                                              • Orleans
                                                              • Atom Editor
                                                              • Automerge
                                                              • Martin Klepman
                                                              • Akka
                                                              • Delta CRDTs
                                                              • Antidote DB
                                                              • Kops
                                                              • Eventual Consistency
                                                              • Causal Consistency
                                                              • ACID Transactions
                                                              • Joe Hellerstein
                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                Support Data Engineering Podcast

                                                                46 min
                                                              • Citus Data: Distributed PostGreSQL for Big Data with Ozgun Erdogan and Craig Kerstiens - Episode 13
                                                                Summary

                                                                PostGreSQL has become one of the most popular and widely used databases, and for good reason. The level of extensibility that it supports has allowed it to be used in virtually every environment. At Citus Data they have built an extension to support running it in a distributed fashion across large volumes of data with parallelized queries for improved performance. In this episode Ozgun Erdogan, the CTO of Citus, and Craig Kerstiens, Citus Product Manager, discuss how the company got started, the work that they are doing to scale out PostGreSQL, and how you can start using it in your environment.

                                                                Preamble
                                                                • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                • Your host is Tobias Macey and today I’m interviewing Ozgun Erdogan and Craig Kerstiens about Citus, worry free PostGreSQL
                                                                • Interview
                                                                  • Introduction
                                                                  • How did you get involved in the area of data management?
                                                                  • Can you describe what Citus is and how the project got started?
                                                                  • Why did you start with Postgres vs. building something from the ground up?
                                                                  • What was the reasoning behind converting Citus from a fork of PostGres to being an extension and releasing an open source version?
                                                                  • How well does Citus work with other Postgres extensions, such as PostGIS, PipelineDB, or Timescale?
                                                                  • How does Citus compare to options such as PostGres-XL or the Postgres compatible Aurora service from Amazon?
                                                                  • How does Citus operate under the covers to enable clustering and replication across multiple hosts?
                                                                  • What are the failure modes of Citus and how does it handle loss of nodes in the cluster?
                                                                  • For someone who is interested in migrating to Citus, what is involved in getting it deployed and moving the data out of an existing system?
                                                                  • How do the different options for leveraging Citus compare to each other and how do you determine which features to release or withhold in the open source version?
                                                                  • Are there any use cases that Citus enables which would be impractical to attempt in native Postgres?
                                                                  • What have been some of the most challenging aspects of building the Citus extension?
                                                                  • What are the situations where you would advise against using Citus?
                                                                  • What are some of the most interesting or impressive uses of Citus that you have seen?
                                                                  • What are some of the features that you have planned for future releases of Citus?
                                                                  • Contact Info
                                                                    • Citus Data
                                                                      • citusdata.com
                                                                      • @citusdata on Twitter
                                                                      • citusdata on GitHub

                                                                      • Craig

                                                                        • Email
                                                                        • Website
                                                                        • @craigkerstiens on Twitter

                                                                        • Ozgun

                                                                          • Email
                                                                          • ozgune on GitHub

                                                                          • Parting Question
                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                            • Links
                                                                              • Citus Data
                                                                              • PostGreSQL
                                                                              • NoSQL
                                                                              • Timescale SQL blog post
                                                                              • PostGIS
                                                                              • PostGreSQL Graph Database
                                                                              • JSONB Data Type
                                                                              • PipelineDB
                                                                              • Timescale
                                                                              • PostGres-XL
                                                                              • Aurora PostGres
                                                                              • Amazon RDS
                                                                              • Streaming Replication
                                                                              • CitusMX
                                                                              • CTE (Common Table Expression)
                                                                              • HipMunk Citus Sharding Blog Post
                                                                              • Wal-e
                                                                              • Wal-g
                                                                              • Heap Analytics
                                                                              • HyperLogLog
                                                                              • C-Store
                                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                Support Data Engineering Podcast

                                                                                47 min
                                                                              • Wallaroo with Sean T. Allen - Episode 12
                                                                                Summary

                                                                                Data oriented applications that need to operate on large, fast-moving sterams of information can be difficult to build and scale due to the need to manage their state. In this episode Sean T. Allen, VP of engineering for Wallaroo Labs, explains how Wallaroo was designed and built to reduce the cognitive overhead of building this style of project. He explains the motivation for building Wallaroo, how it is implemented, and how you can start using it today.

                                                                                Preamble
                                                                                • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                                • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                • Your host is Tobias Macey and today I’m interviewing Sean T. Allen about Wallaroo, a framework for building and operating stateful data applications at scale
                                                                                • Interview
                                                                                  • Introduction
                                                                                  • How did you get involved in the area of data engineering?
                                                                                  • What is Wallaroo and how did the project get started?
                                                                                  • What is the Pony language, and what features does it have that make it well suited for the problem area that you are focusing on?
                                                                                  • Why did you choose to focus first on Python as the language for interacting with Wallaroo and how is that integration implemented?
                                                                                  • How is Wallaroo architected internally to allow for distributed state management?
                                                                                    • Is the state persistent, or is it only maintained long enough to complete the desired computation?
                                                                                    • If so, what format do you use for long term storage of the data?

                                                                                    • What have been the most challenging aspects of building the Wallaroo platform?

                                                                                    • Which axes of the CAP theorem have you optimized for?

                                                                                    • For someone who wants to build an application on top of Wallaroo, what is involved in getting started?

                                                                                    • Once you have a working application, what resources are necessary for deploying to production and what are the scaling factors?

                                                                                      • What are the failure modes that users of Wallaroo need to account for in their application or infrastructure?

                                                                                      • What are some situations or problem types for which Wallaroo would be the wrong choice?

                                                                                      • What are some of the most interesting or unexpected uses of Wallaroo that you have seen?

                                                                                      • What do you have planned for the future of Wallaroo?

                                                                                      • Contact Info
                                                                                        • IRC
                                                                                        • Mailing List
                                                                                        • Wallaroo Labs Twitter
                                                                                        • Email
                                                                                        • Personal Twitter
                                                                                        • Parting Question
                                                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                          • Links
                                                                                            • Wallaroo Labs
                                                                                            • Storm Applied
                                                                                            • Apache Storm
                                                                                            • Risk Analysis
                                                                                            • Pony Language
                                                                                            • Erlang
                                                                                            • Akka
                                                                                            • Tail Latency
                                                                                            • High Performance Computing
                                                                                            • Python
                                                                                            • Apache Software Foundation
                                                                                            • Beyond Distributed Transactions: An Apostate’s View
                                                                                            • Consistent Hashing
                                                                                            • Jepsen
                                                                                            • Lineage Driven Fault Injection
                                                                                            • Chaos Engineering
                                                                                            • QCon 2016 Talk
                                                                                            • Codemesh in London: How did I get here?
                                                                                            • CAP Theorem
                                                                                            • CRDT
                                                                                            • Sync Free Project
                                                                                            • Basho
                                                                                            • Wallaroo on GitHub
                                                                                            • Docker
                                                                                            • Puppet
                                                                                            • Chef
                                                                                            • Ansible
                                                                                            • SaltStack
                                                                                            • Kafka
                                                                                            • TCP
                                                                                            • Dask
                                                                                            • Data Engineering Episode About Dask
                                                                                            • Beowulf Cluster
                                                                                            • Redis
                                                                                            • Flink
                                                                                            • Haskell
                                                                                            • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                              Support Data Engineering Podcast

                                                                                              1 hr
                                                                                            • SiriDB: Scalable Open Source Timeseries Database with Jeroen van der Heijden - Episode 11
                                                                                              Summary

                                                                                              Time series databases have long been the cornerstone of a robust metrics system, but the existing options are often difficult to manage in production. In this episode Jeroen van der Heijden explains his motivation for writing a new database, SiriDB, the challenges that he faced in doing so, and how it works under the hood.

                                                                                              Preamble
                                                                                              • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                              • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                              • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                                              • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                              • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                              • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                              • Your host is Tobias Macey and today I’m interviewing Jeroen van der Heijden about SiriDB, a next generation time series database
                                                                                              • Interview
                                                                                                • Introduction
                                                                                                • How did you get involved in the area of data engineering?
                                                                                                • What is SiriDB and how did the project get started?
                                                                                                  • What was the inspiration for the name?

                                                                                                  • What was the landscape of time series databases at the time that you first began work on Siri?

                                                                                                  • How does Siri compare to other time series databases such as InfluxDB, Timescale, KairosDB, etc.?

                                                                                                  • What do you view as the competition for Siri?

                                                                                                  • How is the server architected and how has the design evolved over the time that you have been working on it?

                                                                                                  • Can you describe how the clustering mechanism functions?

                                                                                                    • Is it possible to create pools with more than two servers?

                                                                                                    • What are the failure modes for SiriDB and where does it fall on the spectrum for the CAP theorem?

                                                                                                    • In the documentation it mentions needing to specify the retention period for the shards when creating a database. What is the reasoning for that and what happens to the individual metrics as they age beyond that time horizon?

                                                                                                    • One of the common difficulties when using a time series database in an operations context is the need for high cardinality of the metrics. How are metrics identified in Siri and is there any support for tagging?

                                                                                                    • What have been the most challenging aspects of building Siri?

                                                                                                    • In what situations or environments would you advise against using Siri?

                                                                                                    • Contact Info
                                                                                                      • joente on Github
                                                                                                      • LinkedIn
                                                                                                      • Parting Question
                                                                                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                        • Links
                                                                                                          • SiriDB
                                                                                                          • Oversight
                                                                                                          • InfluxDB
                                                                                                          • LevelDB
                                                                                                          • OpenTSDB
                                                                                                          • Timescale DB
                                                                                                          • KairosDB
                                                                                                          • Write Ahead Log
                                                                                                          • Grafana
                                                                                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                            Support Data Engineering Podcast

                                                                                                            34 min
                                                                                                          • Confluent Schema Registry with Ewen Cheslack-Postava - Episode 10
                                                                                                            Summary

                                                                                                            To process your data you need to know what shape it has, which is why schemas are important. When you are processing that data in multiple systems it can be difficult to ensure that they all have an accurate representation of that schema, which is why Confluent has built a schema registry that plugs into Kafka. In this episode Ewen Cheslack-Postava explains what the schema registry is, how it can be used, and how they built it. He also discusses how it can be extended for other deployment targets and use cases, and additional features that are planned for future releases.

                                                                                                            Preamble
                                                                                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                            • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                            • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                                                            • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                            • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                            • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                            • Your host is Tobias Macey and today I’m interviewing Ewen Cheslack-Postava about the Confluent Schema Registry
                                                                                                            • Interview
                                                                                                              • Introduction
                                                                                                              • How did you get involved in the area of data engineering?
                                                                                                              • What is the schema registry and what was the motivating factor for building it?
                                                                                                              • If you are using Avro, what benefits does the schema registry provide over and above the capabilities of Avro’s built in schemas?
                                                                                                              • How did you settle on Avro as the format to support and what would be involved in expanding that support to other serialization options?
                                                                                                              • Conversely, what would be involved in using a storage backend other than Kafka?
                                                                                                              • What are some of the alternative technologies available for people who aren’t using Kafka in their infrastructure?
                                                                                                              • What are some of the biggest challenges that you faced while designing and building the schema registry?
                                                                                                              • What is the tipping point in terms of system scale or complexity when it makes sense to invest in a shared schema registry and what are the alternatives for smaller organizations?
                                                                                                              • What are some of the features or enhancements that you have in mind for future work?
                                                                                                              • Contact Info
                                                                                                                • ewencp on GitHub
                                                                                                                • Website
                                                                                                                • @ewencp on Twitter
                                                                                                                • Parting Question
                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                  • Links
                                                                                                                    • Kafka
                                                                                                                    • Confluent
                                                                                                                    • Schema Registry
                                                                                                                    • Second Life
                                                                                                                    • Eve Online
                                                                                                                    • Yes, Virginia, You Really Do Need a Schema Registry
                                                                                                                    • JSON-Schema
                                                                                                                    • Parquet
                                                                                                                    • Avro
                                                                                                                    • Thrift
                                                                                                                    • Protocol Buffers
                                                                                                                    • Zookeeper
                                                                                                                    • Kafka Connect
                                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                      Support Data Engineering Podcast

                                                                                                                      50 min
                                                                                                                    • data.world with Bryon Jacob - Episode 9
                                                                                                                      Summary

                                                                                                                      We have tools and platforms for collaborating on software projects and linking them together, wouldn’t it be nice to have the same capabilities for data? The team at data.world are working on building a platform to host and share data sets for public and private use that can be linked together to build a semantic web of information. The CTO, Bryon Jacob, discusses how the company got started, their mission, and how they have built and evolved their technical infrastructure.

                                                                                                                      Preamble
                                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                                      • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                                      • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                                                                      • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                                      • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                                      • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                                      • This is your host Tobias Macey and today I’m interviewing Bryon Jacob about the technology and purpose that drive data.world
                                                                                                                      • Interview
                                                                                                                        • Introduction
                                                                                                                        • How did you first get involved in the area of data management?
                                                                                                                        • What is data.world and what is its mission and how does your status as a B Corporation tie into that?
                                                                                                                        • The platform that you have built provides hosting for a large variety of data sizes and types. What does the technical infrastructure consist of and how has that architecture evolved from when you first launched?
                                                                                                                        • What are some of the scaling problems that you have had to deal with as the amount and variety of data that you host has increased?
                                                                                                                        • What are some of the technical challenges that you have been faced with that are unique to the task of hosting a heterogeneous assortment of data sets that intended for shared use?
                                                                                                                        • How do you deal with issues of privacy or compliance associated with data sets that are submitted to the platform?
                                                                                                                        • What are some of the improvements or new capabilities that you are planning to implement as part of the data.world platform?
                                                                                                                        • What are the projects or companies that you consider to be your competitors?
                                                                                                                        • What are some of the most interesting or unexpected uses of the data.world platform that you are aware of?
                                                                                                                        • Contact Information
                                                                                                                          • @bryonjacob on Twitter
                                                                                                                          • bryonjacob on GitHub
                                                                                                                          • LinkedIn
                                                                                                                          • Parting Question
                                                                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                            • Links
                                                                                                                              • data.world
                                                                                                                              • HomeAway
                                                                                                                              • Semantic Web
                                                                                                                              • Knowledge Engineering
                                                                                                                              • Ontology
                                                                                                                              • Open Data
                                                                                                                              • RDF
                                                                                                                              • CSVW
                                                                                                                              • SPARQL
                                                                                                                              • DBPedia
                                                                                                                              • Triplestore
                                                                                                                              • Header Dictionary Triples
                                                                                                                              • Apache Jena
                                                                                                                              • Tabula
                                                                                                                              • Tableau Connector
                                                                                                                              • Excel Connector
                                                                                                                              • Data For Democracy
                                                                                                                              • Jonathan Morgan
                                                                                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                Support Data Engineering Podcast

                                                                                                                                47 min
                                                                                                                              • Data Serialization Formats with Doug Cutting and Julien Le Dem - Episode 8
                                                                                                                                Summary

                                                                                                                                With the wealth of formats for sending and storing data it can be difficult to determine which one to use. In this episode Doug Cutting, creator of Avro, and Julien Le Dem, creator of Parquet, dig into the different classes of serialization formats, what their strengths are, and how to choose one for your workload. They also discuss the role of Arrow as a mechanism for in-memory data sharing and how hardware evolution will influence the state of the art for data formats.

                                                                                                                                Preamble
                                                                                                                                • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                                                • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                                                • Continuous delivery lets you get new features in front of your users as fast as possible without introducing bugs or breaking production and GoCD is the open source platform made by the people at Thoughtworks who wrote the book about it. Go to dataengineeringpodcast.com/gocd to download and launch it today. Enterprise add-ons and professional support are available for added peace of mind.
                                                                                                                                • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                                                • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                                                • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                                                • This is your host Tobias Macey and today I’m interviewing Julien Le Dem and Doug Cutting about data serialization formats and how to pick the right one for your systems.
                                                                                                                                • Interview
                                                                                                                                  • Introduction
                                                                                                                                  • How did you first get involved in the area of data management?
                                                                                                                                  • What are the main serialization formats used for data storage and analysis?
                                                                                                                                  • What are the tradeoffs that are offered by the different formats?
                                                                                                                                  • How have the different storage and analysis tools influenced the types of storage formats that are available?
                                                                                                                                  • You’ve each developed a new on-disk data format, Avro and Parquet respectively. What were your motivations for investing that time and effort?
                                                                                                                                  • Why is it important for data engineers to carefully consider the format in which they transfer their data between systems?
                                                                                                                                    • What are the switching costs involved in moving from one format to another after you have started using it in a production system?
                                                                                                                                    • What are some of the new or upcoming formats that you are each excited about?
                                                                                                                                    • How do you anticipate the evolving hardware, patterns, and tools for processing data to influence the types of storage formats that maintain or grow their popularity?
                                                                                                                                    • Contact Information
                                                                                                                                      • Doug:
                                                                                                                                        • cutting on GitHub
                                                                                                                                        • Blog
                                                                                                                                        • @cutting on Twitter
                                                                                                                                        • Julien
                                                                                                                                          • Email
                                                                                                                                          • @J_ on Twitter
                                                                                                                                          • Blog
                                                                                                                                          • julienledem on GitHub
                                                                                                                                          • Links
                                                                                                                                            • Apache Avro
                                                                                                                                            • Apache Parquet
                                                                                                                                            • Apache Arrow
                                                                                                                                            • Hadoop
                                                                                                                                            • Apache Pig
                                                                                                                                            • Xerox Parc
                                                                                                                                            • Excite
                                                                                                                                            • Nutch
                                                                                                                                            • Vertica
                                                                                                                                            • Dremel White Paper
                                                                                                                                              • Twitter Blog on Release of Parquet
                                                                                                                                              • CSV
                                                                                                                                              • XML
                                                                                                                                              • Hive
                                                                                                                                              • Impala
                                                                                                                                              • Presto
                                                                                                                                              • Spark SQL
                                                                                                                                              • Brotli
                                                                                                                                              • ZStandard
                                                                                                                                              • Apache Drill
                                                                                                                                              • Trevni
                                                                                                                                              • Apache Calcite
                                                                                                                                              • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                Support Data Engineering Podcast

                                                                                                                                                52 min

                                                                                                                                              About Data Engineering Podcast

                                                                                                                                              From the publisher's feed

                                                                                                                                              This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

                                                                                                                                              More shows like Data Engineering Podcast

                                                                                                                                              This Week in Startups by Jason Calacanis

                                                                                                                                              This Week in Startups

                                                                                                                                              1,289 Listeners

                                                                                                                                              The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                              The Changelog: Software Development, Open Source

                                                                                                                                              286 Listeners

                                                                                                                                              The a16z Show by Andreessen Horowitz

                                                                                                                                              The a16z Show

                                                                                                                                              1,089 Listeners

                                                                                                                                              Software Engineering Daily by Software Engineering Daily

                                                                                                                                              Software Engineering Daily

                                                                                                                                              622 Listeners

                                                                                                                                              Risky Business by Risky Business Media

                                                                                                                                              Risky Business

                                                                                                                                              374 Listeners

                                                                                                                                              Talk Python To Me by Michael Kennedy

                                                                                                                                              Talk Python To Me

                                                                                                                                              582 Listeners

                                                                                                                                              Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                              Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                              304 Listeners

                                                                                                                                              NVIDIA AI Podcast by NVIDIA

                                                                                                                                              NVIDIA AI Podcast

                                                                                                                                              337 Listeners

                                                                                                                                              Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                              Syntax - Tasty Web Development Treats

                                                                                                                                              985 Listeners

                                                                                                                                              Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                              Practical AI

                                                                                                                                              203 Listeners

                                                                                                                                              Dwarkesh Podcast by Dwarkesh Patel

                                                                                                                                              Dwarkesh Podcast

                                                                                                                                              565 Listeners

                                                                                                                                              The Data Engineering Show by The Firebolt Data Bros

                                                                                                                                              The Data Engineering Show

                                                                                                                                              8 Listeners

                                                                                                                                              Latent Space: The AI Engineer Podcast by Latent.Space

                                                                                                                                              Latent Space: The AI Engineer Podcast

                                                                                                                                              102 Listeners

                                                                                                                                              This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                                                                                                                                              This Day in AI Podcast

                                                                                                                                              222 Listeners

                                                                                                                                              The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                                                                                                                                              The AI Daily Brief: Artificial Intelligence News and Analysis

                                                                                                                                              685 Listeners