Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Data Engineering Weekly with Joe Crobak - Episode 27
    Summary

    The rate of change in the data engineering industry is alternately exciting and exhausting. Joe Crobak found his way into the work of data management by accident as so many of us do. After being engrossed with researching the details of distributed systems and big data management for his work he began sharing his findings with friends. This led to his creation of the Hadoop Weekly newsletter, which he recently rebranded as the Data Engineering Weekly newsletter. In this episode he discusses his experiences working as a data engineer in industry and at the USDS, his motivations and methods for creating a newsleteter, and the insights that he has gleaned from it.

    Preamble
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
    • Your host is Tobias Macey and today I’m interviewing Joe Crobak about his work maintaining the Data Engineering Weekly newsletter, and the challenges of keeping up with the data engineering industry.
    • Interview
      • Introduction
      • How did you get involved in the area of data management?
      • What are some of the projects that you have been involved in that were most personally fulfilling?
        • As an engineer at the USDS working on the healthcare.gov and medicare systems, what were some of the approaches that you used to manage sensitive data?
        • Healthcare.gov has a storied history, how did the systems for processing and managing the data get architected to handle the amount of load that it was subjected to?

        • What was your motivation for starting a newsletter about the Hadoop space?

          • Can you speak to your reasoning for the recent rebranding of the newsletter?

          • How much of the content that you surface in your newsletter is found during your day-to-day work, versus explicitly searching for it?

          • After over 5 years of following the trends in data analytics and data infrastructure what are some of the most interesting or surprising developments?

            • What have you found to be the fundamental skills or areas of experience that have maintained relevance as new technologies in data engineering have emerged?

            • What is your workflow for finding and curating the content that goes into your newsletter?

            • What is your personal algorithm for filtering which articles, tools, or commentary gets added to the final newsletter?

            • How has your experience managing the newsletter influenced your areas of focus in your work and vice-versa?

            • What are your plans going forward?

            • Contact Info
              • Data Eng Weekly
              • Email
              • Twitter – @joecrobak
              • Twitter – @dataengweekly
              • Parting Question
                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                • Links
                  • USDS
                  • National Labs
                  • Cray
                  • Amazon EMR (Elastic Map-Reduce)
                  • Recommendation Engine
                  • Netflix Prize
                  • Hadoop
                  • Cloudera
                  • Puppet
                  • healthcare.gov
                  • Medicare
                  • Quality Payment Program
                  • HIPAA
                  • NIST National Institute of Standards and Technology
                  • PII (Personally Identifiable Information)
                  • Threat Modeling
                  • Apache JBoss
                  • Apache Web Server
                  • MarkLogic
                  • JMS (Java Message Service)
                  • Load Balancer
                  • COBOL
                  • Hadoop Weekly
                  • Data Engineering Weekly
                  • Foursquare
                  • NiFi
                  • Kubernetes
                  • Spark
                  • Flink
                  • Stream Processing
                  • DataStax
                  • RSS
                  • The Flavors of Data Science and Engineering
                  • CQRS
                  • Change Data Capture
                  • Jay Kreps
                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                    Support Data Engineering Podcast

                    44 min
                  • Defining DataOps with Chris Bergh - Episode 26
                    Summary

                    Managing an analytics project can be difficult due to the number of systems involved and the need to ensure that new information can be delivered quickly and reliably. That challenge can be met by adopting practices and principles from lean manufacturing and agile software development, and the cross-functional collaboration, feedback loops, and focus on automation in the DevOps movement. In this episode Christopher Bergh discusses ways that you can start adding reliability and speed to your workflow to deliver results with confidence and consistency.

                    Preamble
                    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                    • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                    • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
                    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                    • Your host is Tobias Macey and today I’m interviewing Christopher Bergh about DataKitchen and the rise of DataOps
                    • Interview
                      • Introduction
                      • How did you get involved in the area of data management?
                      • How do you define DataOps?
                        • How does it compare to the practices encouraged by the DevOps movement?
                        • How does it relate to or influence the role of a data engineer?

                        • How does a DataOps oriented workflow differ from other existing approaches for building data platforms?

                        • One of the aspects of DataOps that you call out is the practice of providing multiple environments to provide a platform for testing the various aspects of the analytics workflow in a non-production context. What are some of the techniques that are available for managing data in appropriate volumes across those deployments?

                        • The practice of testing logic as code is fairly well understood and has a large set of existing tools. What have you found to be some of the most effective methods for testing data as it flows through a system?

                        • One of the practices of DevOps is to create feedback loops that can be used to ensure that business needs are being met. What are the metrics that you track in your platform to define the value that is being created and how the various steps in the workflow are proceeding toward that goal?

                          • In order to keep feedback loops fast it is necessary for tests to run quickly. How do you balance the need for larger quantities of data to be used for verifying scalability/performance against optimizing for cost and speed in non-production environments?

                          • How does the DataKitchen platform simplify the process of operationalizing a data analytics workflow?

                          • As the need for rapid iteration and deployment of systems to capture, store, process, and analyze data becomes more prevalent how do you foresee that feeding back into the ways that the landscape of data tools are designed and developed?

                          • Contact Info
                            • LinkedIn
                            • @ChrisBergh on Twitter
                            • Email
                            • Parting Question
                              • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                              • Links
                                • DataOps Manifesto
                                • DataKitchen
                                • 2017: The Year Of DataOps
                                • Air Traffic Control
                                • Chief Data Officer (CDO)
                                • Gartner
                                • W. Edwards Deming
                                • DevOps
                                • Total Quality Management (TQM)
                                • Informatica
                                • Talend
                                • Agile Development
                                • Cattle Not Pets
                                • IDE (Integrated Development Environment)
                                • Tableau
                                • Delphix
                                • Dremio
                                • Pachyderm
                                • Continuous Delivery by Jez Humble and Dave Farley
                                • SLAs (Service Level Agreements)
                                • XKCD Image Recognition Comic
                                • Airflow
                                • Luigi
                                • DataKitchen Documentation
                                • Continuous Integration
                                • Continous Delivery
                                • Docker
                                • Version Control
                                • Git
                                • Looker
                                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                  Support Data Engineering Podcast

                                  55 min
                                • ThreatStack: Data Driven Cloud Security with Pete Cheslock and Patrick Cable - Episode 25
                                  Summary

                                  Cloud computing and ubiquitous virtualization have changed the ways that our applications are built and deployed. This new environment requires a new way of tracking and addressing the security of our systems. ThreatStack is a platform that collects all of the data that your servers generate and monitors for unexpected anomalies in behavior that would indicate a breach and notifies you in near-realtime. In this episode ThreatStack’s director of operations, Pete Cheslock, and senior infrastructure security engineer, Patrick Cable, discuss the data infrastructure that supports their platform, how they capture and process the data from client systems, and how that information can be used to keep your systems safe from attackers.

                                  Preamble
                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                  • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                  • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
                                  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                  • Your host is Tobias Macey and today I’m interviewing Pete Cheslock and Pat Cable about the data infrastructure and security controls at ThreatStack
                                  • Interview
                                    • Introduction
                                    • How did you get involved in the area of data management?
                                    • Why don’t you start by explaining what ThreatStack does?
                                      • What was lacking in the existing options (services and self-hosted/open source) that ThreatStack solves for?

                                      • Can you describe the type(s) of data that you collect and how it is structured?

                                      • What is the high level data infrastructure that you use for ingesting, storing, and analyzing your customer data?

                                        • How do you ensure a consistent format of the information that you receive?
                                        • How do you ensure that the various pieces of your platform are deployed using the proper configurations and operating as intended?
                                        • How much configuration do you provide to the end user in terms of the captured data, such as sampling rate or additional context?

                                        • I understand that your original architecture used RabbitMQ as your ingest mechanism, which you then migrated to Kafka. What was your initial motivation for that change?

                                          • How much of a benefit has that been in terms of overall complexity and cost (both time and infrastructure)?

                                          • How do you ensure the security and provenance of the data that you collect as it traverses your infrastructure?

                                          • What are some of the most common vulnerabilities that you detect in your client’s infrastructure?

                                          • For someone who wants to start using ThreatStack, what does the setup process look like?

                                          • What have you found to be the most challenging aspects of building and managing the data processes in your environment?

                                          • What are some of the projects that you have planned to improve the capacity or capabilities of your infrastructure?

                                          • Contact Info
                                            • Pete Cheslock
                                              • @petecheslock on Twitter
                                              • Website
                                              • petecheslock on GitHub

                                              • Patrick Cable

                                                • @patcable on Twitter
                                                • Website
                                                • patcable on GitHub

                                                • ThreatStack

                                                  • Website
                                                  • @threatstack on Twitter
                                                  • threatstack on GitHub

                                                  • Parting Question
                                                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                    • Links
                                                      • ThreatStack
                                                      • SecDevOps
                                                      • Sonian
                                                      • EC2
                                                      • Snort
                                                      • Snorby
                                                      • Suricata
                                                      • Tripwire
                                                      • Syscall (System Call)
                                                      • AuditD
                                                      • CloudTrail
                                                      • Naxsi
                                                      • Cloud Native
                                                      • File Integrity Monitoring (FIM)
                                                      • Amazon Web Services (AWS)
                                                      • RabbitMQ
                                                      • ZeroMQ
                                                      • Kafka
                                                      • Spark
                                                      • Slack
                                                      • PagerDuty
                                                      • JSON
                                                      • Microservices
                                                      • Cassandra
                                                      • ElasticSearch
                                                      • Sensu
                                                      • Service Discovery
                                                      • Honeypot
                                                      • Kubernetes
                                                      • PostGreSQL
                                                      • Druid
                                                      • Flink
                                                      • Launch Darkly
                                                      • Chef
                                                      • Consul
                                                      • Terraform
                                                      • CloudFormation
                                                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                        Support Data Engineering Podcast

                                                        52 min
                                                      • MarketStore: Managing Timeseries Financial Data with Hitoshi Harada and Christopher Ryan - Episode 24
                                                        Summary

                                                        The data that is used in financial markets is time oriented and multidimensional, which makes it difficult to manage in either relational or timeseries databases. To make this information more manageable the team at Alapaca built a new data store specifically for retrieving and analyzing data generated by trading markets. In this episode Hitoshi Harada, the CTO of Alapaca, and Christopher Ryan, their lead software engineer, explain their motivation for building MarketStore, how it operates, and how it has helped to simplify their development workflows.

                                                        Preamble
                                                        • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                        • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                        • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
                                                        • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                        • Your host is Tobias Macey and today I’m interviewing Christopher Ryan and Hitoshi Harada about MarketStore, a storage server for large volumes of financial timeseries data
                                                        • Interview
                                                          • Introduction
                                                          • How did you get involved in the area of data management?
                                                          • What was your motivation for creating MarketStore?
                                                          • What are the characteristics of financial time series data that make it challenging to manage?
                                                          • What are some of the workflows that MarketStore is used for at Alpaca and how were they managed before it was available?
                                                          • With MarketStore’s data coming from multiple third party services, how are you managing to keep the DB up-to-date and in sync with those services?
                                                            • What is the worst case scenario if there is a total failure in the data store?
                                                            • What guards have you built to prevent such a situation from occurring?

                                                            • Since MarketStore is used for querying and analyzing data having to do with financial markets and there are potentially large quantities of money being staked on the results of that analysis, how do you ensure that the operations being performed in MarketStore are accurate and repeatable?

                                                            • What were the most challenging aspects of building MarketStore and integrating it into the rest of your systems?

                                                            • Motivation for open sourcing the code?

                                                            • What is the next planned major feature for MarketStore, and what use-case is it aiming to support?

                                                            • Contact Info
                                                              • Christopher
                                                                • Email

                                                                • Hitoshi

                                                                  • Email

                                                                  • Parting Question
                                                                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                    • Links
                                                                      • MarketStore
                                                                        • GitHub
                                                                        • Release Announcement

                                                                        • Alpaca

                                                                        • IBM

                                                                        • DB2

                                                                        • GreenPlum

                                                                        • Algorithmic Trading

                                                                        • Backtesting

                                                                        • OHLC (Open-High-Low-Close)

                                                                        • HDF5

                                                                        • Golang

                                                                        • C++

                                                                        • Timeseries Database List

                                                                        • InfluxDB

                                                                        • JSONRPC

                                                                        • Slait

                                                                        • CircleCI

                                                                        • GDAX

                                                                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                          Support Data Engineering Podcast

                                                                          34 min
                                                                        • Stretching The Elastic Stack with Philipp Krenn - Episode 23
                                                                          Summary

                                                                          Search is a common requirement for applications of all varieties. Elasticsearch was built to make it easy to include search functionality in projects built in any language. From that foundation, the rest of the Elastic Stack has been built, expanding to many more use cases in the proces. In this episode Philipp Krenn describes the various pieces of the stack, how they fit together, and how you can use them in your infrastructure to store, search, and analyze your data.

                                                                          Preamble
                                                                          • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                          • When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
                                                                          • For complete visibility into the health of your pipeline, including deployment tracking, and powerful alerting driven by machine-learning, DataDog has got you covered. With their monitoring, metrics, and log collection agent, including extensive integrations and distributed tracing, you’ll have everything you need to find and fix performance bottlenecks in no time. Go to dataengineeringpodcast.com/datadog today to start your free 14 day trial and get a sweet new T-Shirt.
                                                                          • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                          • Your host is Tobias Macey and today I’m interviewing Philipp Krenn about the Elastic Stack and the ways that you can use it in your systems
                                                                          • Interview
                                                                            • Introduction
                                                                            • How did you get involved in the area of data management?
                                                                            • The Elasticsearch product has been around for a long time and is widely known, but can you give a brief overview of the other components that make up the Elastic Stack and how they work together?
                                                                            • Beyond the common pattern of using Elasticsearch as a search engine connected to a web application, what are some of the other use cases for the various pieces of the stack?
                                                                            • What are the common scaling bottlenecks that users should be aware of when they are dealing with large volumes of data?
                                                                            • What do you consider to be the biggest competition to the Elastic Stack as you expand the capabilities and target usage patterns?
                                                                            • What are the biggest challenges that you are tackling in the Elastic stack, technical or otherwise?
                                                                            • What are the biggest challenges facing Elastic as a company in the near to medium term?
                                                                            • Open source as a business model: https://www.elastic.co/blog/doubling-down-on-open?utm_source=rss&utm_medium=rss
                                                                            • What is the vision for Elastic and the Elastic Stack going forward and what new features or functionality can we look forward to?
                                                                            • Contact Info
                                                                              • @xeraa on Twitter
                                                                              • xeraa on GitHub
                                                                              • Website
                                                                              • Email
                                                                              • Parting Question
                                                                                • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                • Links
                                                                                  • Elastic
                                                                                  • Vienna – Capital of Austria
                                                                                  • What Is Developer Advocacy?
                                                                                  • NoSQL
                                                                                  • MongoDB
                                                                                  • Elasticsearch
                                                                                  • Cassandra
                                                                                  • Neo4J
                                                                                  • Hazelcast
                                                                                  • Apache Lucene
                                                                                  • Logstash
                                                                                  • Kibana
                                                                                  • Beats
                                                                                  • X-Pack
                                                                                  • ELK Stack
                                                                                  • Metrics
                                                                                  • APM (Application Performance Monitoring)
                                                                                  • GeoJSON
                                                                                  • Split Brain
                                                                                  • Elasticsearch Ingest Nodes
                                                                                  • PacketBeat
                                                                                  • Elastic Cloud
                                                                                  • Elasticon
                                                                                  • Kibana Canvas
                                                                                  • SwiftType
                                                                                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                    Support Data Engineering Podcast

                                                                                    52 min
                                                                                  • Database Refactoring Patterns with Pramod Sadalage - Episode 22
                                                                                    Summary

                                                                                    As software lifecycles move faster, the database needs to be able to keep up. Practices such as version controlled migration scripts and iterative schema evolution provide the necessary mechanisms to ensure that your data layer is as agile as your application. Pramod Sadalage saw the need for these capabilities during the early days of the introduction of modern development practices and co-authored a book to codify a large number of patterns to aid practitioners, and in this episode he reflects on the current state of affairs and how things have changed over the past 12 years.

                                                                                    Preamble
                                                                                    • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                    • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                    • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                    • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                    • Your host is Tobias Macey and today I’m interviewing Pramod Sadalage about refactoring databases and integrating database design into an iterative development workflow
                                                                                    • Interview
                                                                                      • Introduction
                                                                                      • How did you get involved in the area of data management?
                                                                                      • You first co-authored Refactoring Databases in 2006. What was the state of software and database system development at the time and why did you find it necessary to write a book on this subject?
                                                                                      • What are the characteristics of a database that make them more difficult to manage in an iterative context?
                                                                                      • How does the practice of refactoring in the context of a database compare to that of software?
                                                                                      • How has the prevalence of data abstractions such as ORMs or ODMs impacted the practice of schema design and evolution?
                                                                                      • Is there a difference in strategy when refactoring the data layer of a system when using a non-relational storage system?
                                                                                      • How has the DevOps movement and the increased focus on automation affected the state of the art in database versioning and evolution?
                                                                                      • What have you found to be the most problematic aspects of databases when trying to evolve the functionality of a system?
                                                                                      • Looking back over the past 12 years, what has changed in the areas of database design and evolution?
                                                                                        • How has the landscape of tooling for managing and applying database versioning changed since you first wrote Refactoring Databases?
                                                                                        • What do you see as the biggest challenges facing us over the next few years?

                                                                                        • Contact Info
                                                                                          • Website
                                                                                          • pramodsadalage on GitHub
                                                                                          • @pramodsadalage on Twitter
                                                                                          • Parting Question
                                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                            • Links
                                                                                              • Database Refactoring
                                                                                                • Website
                                                                                                • Book

                                                                                                • Thoughtworks

                                                                                                • Martin Fowler

                                                                                                • Agile Software Development

                                                                                                • XP (Extreme Programming)

                                                                                                • Continuous Integration

                                                                                                  • The Book
                                                                                                  • Wikipedia

                                                                                                  • Test First Development

                                                                                                  • DDL (Data Definition Language)

                                                                                                  • DML (Data Modification Language)

                                                                                                  • DevOps

                                                                                                  • Flyway

                                                                                                  • Liquibase

                                                                                                  • DBMaintain

                                                                                                  • Hibernate

                                                                                                  • SQLAlchemy

                                                                                                  • ORM (Object Relational Mapper)

                                                                                                  • ODM (Object Document Mapper)

                                                                                                  • NoSQL

                                                                                                  • Document Database

                                                                                                  • MongoDB

                                                                                                  • OrientDB

                                                                                                  • CouchBase

                                                                                                  • CassandraDB

                                                                                                  • Neo4j

                                                                                                  • ArangoDB

                                                                                                  • Unit Testing

                                                                                                  • Integration Testing

                                                                                                  • OLAP (On-Line Analytical Processing)

                                                                                                  • OLTP (On-Line Transaction Processing)

                                                                                                  • Data Warehouse

                                                                                                  • Docker

                                                                                                  • QA==Quality Assurance

                                                                                                  • HIPAA (Health Insurance Portability and Accountability Act)

                                                                                                  • PCI DSS (Payment Card Industry Data Security Standard)

                                                                                                  • Polyglot Persistence

                                                                                                  • Toplink Java ORM

                                                                                                  • Ruby on Rails

                                                                                                  • ActiveRecord Gem

                                                                                                  • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                    Support Data Engineering Podcast

                                                                                                    50 min
                                                                                                  • The Future Data Economy with Roger Chen - Episode 21
                                                                                                    Summary

                                                                                                    Data is an increasingly sought after raw material for business in the modern economy. One of the factors driving this trend is the increase in applications for machine learning and AI which require large quantities of information to work from. As the demand for data becomes more widespread the market for providing it will begin transform the ways that information is collected and shared among and between organizations. With his experience as a chair for the O’Reilly AI conference and an investor for data driven businesses Roger Chen is well versed in the challenges and solutions being facing us. In this episode he shares his perspective on the ways that businesses can work together to create shared data resources that will allow them to reduce the redundancy of their foundational data and improve their overall effectiveness in collecting useful training sets for their particular products.

                                                                                                    Preamble
                                                                                                    • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                    • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                    • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                    • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                    • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                    • A few announcements:
                                                                                                      • The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
                                                                                                      • If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.

                                                                                                      • Your host is Tobias Macey and today I’m interviewing Roger Chen about data liquidity and its impact on our future economies

                                                                                                      • Interview
                                                                                                        • Introduction
                                                                                                        • How did you get involved in the area of data management?
                                                                                                        • You wrote an essay discussing how the increasing usage of machine learning and artificial intelligence applications will result in a demand for data that necessitates what you refer to as ‘Data Liquidity’. Can you explain what you mean by that term?
                                                                                                        • What are some examples of the types of data that you envision as being foundational to multiple organizations and problem domains?
                                                                                                        • Can you provide some examples of the structures that could be created to facilitate data sharing across organizational boundaries?
                                                                                                        • Many companies view their data as a strategic asset and are therefore loathe to provide access to other individuals or organizations. What encouragement can you provide that would convince them to externalize any of that information?
                                                                                                        • What kinds of storage and transmission infrastructure and tooling are necessary to allow for wider distribution of, and collaboration on, data assets?
                                                                                                        • What do you view as being the privacy implications from creating and sharing these larger pools of data inventory?
                                                                                                        • What do you view as some of the technical challenges associated with identifying and separating shared data from those that are specific to the business model of the organization?
                                                                                                        • With broader access to large data sets, how do you anticipate that impacting the types of businesses or products that are possible for smaller organizations?
                                                                                                        • Contact Info
                                                                                                          • @rgrchen on Twitter
                                                                                                          • LinkedIn
                                                                                                          • Angel List
                                                                                                          • Parting Question
                                                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                            • Links
                                                                                                              • Electrical Engineering
                                                                                                              • Berkeley
                                                                                                              • Silicon Nanophotonics
                                                                                                              • Data Liquidity In The Age Of Inference
                                                                                                              • Data Silos
                                                                                                              • Example of a Data Commons Cooperative
                                                                                                              • Google Maps Moat: An article describing how Google Maps has refined raw data to create a new product
                                                                                                              • Genomics
                                                                                                              • Phenomics
                                                                                                              • ImageNet
                                                                                                              • Open Data
                                                                                                              • Data Brokerage
                                                                                                              • Smart Contracts
                                                                                                              • IPFS
                                                                                                              • Dat Protocol
                                                                                                              • Homomorphic Encryption
                                                                                                              • FileCoin
                                                                                                              • Data Programming
                                                                                                              • Snorkel
                                                                                                                • Website
                                                                                                                • Podcast Interview

                                                                                                                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                  Support Data Engineering Podcast

                                                                                                                  43 min
                                                                                                                • Honeycomb Data Infrastructure with Sam Stokes - Episode 20
                                                                                                                  Summary

                                                                                                                  One of the sources of data that often gets overlooked is the systems that we use to run our businesses. This data is not used to directly provide value to customers or understand the functioning of the business, but it is still a critical component of a successful system. Sam Stokes is an engineer at Honeycomb where he helps to build a platform that is able to capture all of the events and context that occur in our production environments and use them to answer all of your questions about what is happening in your system right now. In this episode he discusses the challenges inherent in capturing and analyzing event data, the tools that his team is using to make it possible, and how this type of knowledge can be used to improve your critical infrastructure.

                                                                                                                  Preamble
                                                                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                                  • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                                  • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                                  • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                                  • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                                  • A few announcements:
                                                                                                                    • There is still time to register for the O’Reilly Strata Conference in San Jose, CA March 5th-8th. Use the link dataengineeringpodcast.com/strata-san-jose to register and save 20%
                                                                                                                    • The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
                                                                                                                    • If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.

                                                                                                                    • Your host is Tobias Macey and today I’m interviewing Sam Stokes about his work at Honeycomb, a modern platform for observability of software systems

                                                                                                                    • Interview
                                                                                                                      • Introduction
                                                                                                                      • How did you get involved in the area of data management?
                                                                                                                      • What is Honeycomb and how did you get started at the company?
                                                                                                                      • Can you start by giving an overview of your data infrastructure and the path that an event takes from ingest to graph?
                                                                                                                      • What are the characteristics of the event data that you are dealing with and what challenges does it pose in terms of processing it at scale?
                                                                                                                      • In addition to the complexities of ingesting and storing data with a high degree of cardinality, being able to quickly analyze it for customer reporting poses a number of difficulties. Can you explain how you have built your systems to facilitate highly interactive usage patterns?
                                                                                                                      • A high degree of visibility into a running system is desirable for developers and systems adminstrators, but they are not always willing or able to invest the effort to fully instrument the code or servers that they want to track. What have you found to be the most difficult aspects of data collection, and do you have any tooling to simplify the implementation for user?
                                                                                                                      • How does Honeycomb compare to other systems that are available off the shelf or as a service, and when is it not the right tool?
                                                                                                                      • What have been some of the most challenging aspects of building, scaling, and marketing Honeycomb?
                                                                                                                      • Contact Info
                                                                                                                        • @samstokes on Twitter
                                                                                                                        • Blog
                                                                                                                        • samstokes on GitHub
                                                                                                                        • Parting Question
                                                                                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                          • Links
                                                                                                                            • Honeycomb
                                                                                                                            • Retriever
                                                                                                                            • Monitoring and Observability
                                                                                                                            • Kafka
                                                                                                                            • Column Oriented Storage
                                                                                                                            • Elasticsearch
                                                                                                                            • Elastic Stack
                                                                                                                            • Django
                                                                                                                            • Ruby on Rails
                                                                                                                            • Heroku
                                                                                                                            • Kubernetes
                                                                                                                            • Launch Darkly
                                                                                                                            • Splunk
                                                                                                                            • Datadog
                                                                                                                            • Cynefin Framework
                                                                                                                            • Go-Lang
                                                                                                                            • Terraform
                                                                                                                            • AWS
                                                                                                                            • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                              Support Data Engineering Podcast

                                                                                                                              42 min
                                                                                                                            • Data Teams with Will McGinnis - Episode 19
                                                                                                                              Summary

                                                                                                                              The responsibilities of a data scientist and a data engineer often overlap and occasionally come to cross purposes. Despite these challenges it is possible for the two roles to work together effectively and produce valuable business outcomes. In this episode Will McGinnis discusses the opinions that he has gained from experience on how data teams can play to their strengths to the benefit of all.

                                                                                                                              Preamble
                                                                                                                              • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                                              • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                                              • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                                              • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                                              • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                                              • A few announcements:
                                                                                                                                • There is still time to register for the O’Reilly Strata Conference in San Jose, CA March 5th-8th. Use the link dataengineeringpodcast.com/strata-san-jose to register and save 20%
                                                                                                                                • The O’Reilly AI Conference is also coming up. Happening April 29th to the 30th in New York it will give you a solid understanding of the latest breakthroughs and best practices in AI for business. Go to dataengineeringpodcast.com/aicon-new-york to register and save 20%
                                                                                                                                • If you work with data or want to learn more about how the projects you have heard about on the show get used in the real world then join me at the Open Data Science Conference in Boston from May 1st through the 4th. It has become one of the largest events for data scientists, data engineers, and data driven businesses to get together and learn how to be more effective. To save 60% off your tickets go to dataengineeringpodcast.com/odsc-east-2018 and register.

                                                                                                                                • Your host is Tobias Macey and today I’m interviewing Will McGinnis about the relationship and boundaries between data engineers and data scientists

                                                                                                                                • Interview
                                                                                                                                  • Introduction
                                                                                                                                  • How did you get involved in the area of data management?
                                                                                                                                  • The terms “Data Scientist” and “Data Engineer” are fluid and seem to have a different meaning for everyone who uses them. Can you share how you define those terms?
                                                                                                                                  • What parallels do you see between the relationships of data engineers and data scientists and those of developers and systems administrators?
                                                                                                                                  • Is there a particular size of organization or problem that serves as a tipping point for when you start to separate the two roles into the responsibilities of more than one person or team?
                                                                                                                                  • What are the benefits of splitting the responsibilities of data engineering and data science?
                                                                                                                                    • What are the disadvantages?

                                                                                                                                    • What are some strategies to ensure successful interaction between data engineers and data scientists?

                                                                                                                                    • How do you view these roles evolving as they become more prevalent across companies and industries?

                                                                                                                                    • Contact Info
                                                                                                                                      • Website
                                                                                                                                      • wdm0006 on GitHub
                                                                                                                                      • @willmcginniser on Twitter
                                                                                                                                      • LinkedIn
                                                                                                                                      • Parting Question
                                                                                                                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                        • Links
                                                                                                                                          • Blog Post: Tendencies of Data Engineers and Data Scientists
                                                                                                                                          • Predikto
                                                                                                                                          • Categorical Encoders
                                                                                                                                          • DevOps
                                                                                                                                          • SciKit-Learn
                                                                                                                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                            Support Data Engineering Podcast

                                                                                                                                            29 min
                                                                                                                                          • TimescaleDB: Fast And Scalable Timeseries with Ajay Kulkarni and Mike Freedman - Episode 18
                                                                                                                                            Summary

                                                                                                                                            As communications between machines become more commonplace the need to store the generated data in a time-oriented manner increases. The market for timeseries data stores has many contenders, but they are not all built to solve the same problems or to scale in the same manner. In this episode the founders of TimescaleDB, Ajay Kulkarni and Mike Freedman, discuss how Timescale was started, the problems that it solves, and how it works under the covers. They also explain how you can start using it in your infrastructure and their plans for the future.

                                                                                                                                            Preamble
                                                                                                                                            • Hello and welcome to the Data Engineering Podcast, the show about modern data infrastructure
                                                                                                                                            • When you’re ready to launch your next project you’ll need somewhere to deploy it. Check out Linode at dataengineeringpodcast.com/linode and get a $20 credit to try out their fast and reliable Linux virtual servers for running your data pipelines or trying out the tools you hear about on the show.
                                                                                                                                            • Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
                                                                                                                                            • You can help support the show by checking out the Patreon page which is linked from the site.
                                                                                                                                            • To help other people find the show you can leave a review on iTunes, or Google Play Music, and tell your friends and co-workers
                                                                                                                                            • Your host is Tobias Macey and today I’m interviewing Ajay Kulkarni and Mike Freedman about Timescale DB, a scalable timeseries database built on top of PostGreSQL
                                                                                                                                            • Interview
                                                                                                                                              • Introduction
                                                                                                                                              • How did you get involved in the area of data management?
                                                                                                                                              • Can you start by explaining what Timescale is and how the project got started?
                                                                                                                                              • The landscape of time series databases is extensive and oftentimes difficult to navigate. How do you view your position in that market and what makes Timescale stand out from the other options?
                                                                                                                                              • In your blog post that explains the design decisions for how Timescale is implemented you call out the fact that the inserted data is largely append only which simplifies the index management. How does Timescale handle out of order timestamps, such as from infrequently connected sensors or mobile devices?
                                                                                                                                              • How is Timescale implemented and how has the internal architecture evolved since you first started working on it?
                                                                                                                                                • What impact has the 10.0 release of PostGreSQL had on the design of the project?
                                                                                                                                                • Is timescale compatible with systems such as Amazon RDS or Google Cloud SQL?

                                                                                                                                                • For someone who wants to start using Timescale what is involved in deploying and maintaining it?

                                                                                                                                                • What are the axes for scaling Timescale and what are the points where that scalability breaks down?

                                                                                                                                                  • Are you aware of anyone who has deployed it on top of Citus for scaling horizontally across instances?

                                                                                                                                                  • What has been the most challenging aspect of building and marketing Timescale?

                                                                                                                                                  • When is Timescale the wrong tool to use for time series data?

                                                                                                                                                  • One of the use cases that you call out on your website is for systems metrics and monitoring. How does Timescale fit into that ecosystem and can it be used along with tools such as Graphite or Prometheus?

                                                                                                                                                  • What are some of the most interesting uses of Timescale that you have seen?

                                                                                                                                                  • Which came first, Timescale the business or Timescale the database, and what is your strategy for ensuring that the open source project and the company around it both maintain their health?

                                                                                                                                                  • What features or improvements do you have planned for future releases of Timescale?

                                                                                                                                                  • Contact Info
                                                                                                                                                    • Ajay
                                                                                                                                                      • LinkedIn
                                                                                                                                                      • @acoustik on Twitter
                                                                                                                                                      • Timescale Blog

                                                                                                                                                      • Mike

                                                                                                                                                        • Website
                                                                                                                                                        • LinkedIn
                                                                                                                                                        • @michaelfreedman on Twitter
                                                                                                                                                        • Timescale Blog

                                                                                                                                                        • Timescale

                                                                                                                                                          • Website
                                                                                                                                                          • @timescaledb on Twitter
                                                                                                                                                          • GitHub

                                                                                                                                                          • Parting Question
                                                                                                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                            • Links
                                                                                                                                                              • Timescale
                                                                                                                                                              • PostGreSQL
                                                                                                                                                              • Citus
                                                                                                                                                              • Timescale Design Blog Post
                                                                                                                                                              • MIT
                                                                                                                                                              • NYU
                                                                                                                                                              • Stanford
                                                                                                                                                              • SDN
                                                                                                                                                              • Princeton
                                                                                                                                                              • Machine Data
                                                                                                                                                              • Timeseries Data
                                                                                                                                                              • List of Timeseries Databases
                                                                                                                                                              • NoSQL
                                                                                                                                                              • Online Transaction Processing (OLTP)
                                                                                                                                                              • Object Relational Mapper (ORM)
                                                                                                                                                              • Grafana
                                                                                                                                                              • Tableau
                                                                                                                                                              • Kafka
                                                                                                                                                              • When Boring Is Awesome
                                                                                                                                                              • PostGreSQL
                                                                                                                                                              • RDS
                                                                                                                                                              • Google Cloud SQL
                                                                                                                                                              • Azure DB
                                                                                                                                                              • Docker
                                                                                                                                                              • Continuous Aggregates
                                                                                                                                                              • Streaming Replication
                                                                                                                                                              • PGPool II
                                                                                                                                                              • Kubernetes
                                                                                                                                                              • Docker Swarm
                                                                                                                                                              • Citus Data
                                                                                                                                                                • Website
                                                                                                                                                                • Data Engineering Podcast Interview

                                                                                                                                                                • Database Indexing

                                                                                                                                                                • B-Tree Index

                                                                                                                                                                • GIN Index

                                                                                                                                                                • GIST Index

                                                                                                                                                                • STE Energy

                                                                                                                                                                • Redis

                                                                                                                                                                • Graphite

                                                                                                                                                                • Prometheus

                                                                                                                                                                • pg_prometheus

                                                                                                                                                                • OpenMetrics Standard Proposal

                                                                                                                                                                • Timescale Parallel Copy

                                                                                                                                                                • Hadoop

                                                                                                                                                                • PostGIS

                                                                                                                                                                • KDB+

                                                                                                                                                                • DevOps

                                                                                                                                                                • Internet of Things

                                                                                                                                                                • MongoDB

                                                                                                                                                                • Elastic

                                                                                                                                                                • DataBricks

                                                                                                                                                                • Apache Spark

                                                                                                                                                                • Confluent

                                                                                                                                                                • New Enterprise Associates

                                                                                                                                                                • MapD

                                                                                                                                                                • Benchmark Ventures

                                                                                                                                                                • Hortonworks

                                                                                                                                                                • 2σ Ventures

                                                                                                                                                                • CockroachDB

                                                                                                                                                                • Cloudflare

                                                                                                                                                                • EMC

                                                                                                                                                                • Timescale Blog: Why SQL is beating NoSQL, and what this means for the future of data

                                                                                                                                                                • The intro and outro music is from

                                                                                                                                                                  1 hr 3 min

                                                                                                                                                                About Data Engineering Podcast

                                                                                                                                                                From the publisher's feed

                                                                                                                                                                This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

                                                                                                                                                                More shows like Data Engineering Podcast

                                                                                                                                                                This Week in Startups by Jason Calacanis

                                                                                                                                                                This Week in Startups

                                                                                                                                                                1,289 Listeners

                                                                                                                                                                The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                                                The Changelog: Software Development, Open Source

                                                                                                                                                                286 Listeners

                                                                                                                                                                The a16z Show by Andreessen Horowitz

                                                                                                                                                                The a16z Show

                                                                                                                                                                1,089 Listeners

                                                                                                                                                                Software Engineering Daily by Software Engineering Daily

                                                                                                                                                                Software Engineering Daily

                                                                                                                                                                622 Listeners

                                                                                                                                                                Risky Business by Risky Business Media

                                                                                                                                                                Risky Business

                                                                                                                                                                374 Listeners

                                                                                                                                                                Talk Python To Me by Michael Kennedy

                                                                                                                                                                Talk Python To Me

                                                                                                                                                                582 Listeners

                                                                                                                                                                Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                                                Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                                                304 Listeners

                                                                                                                                                                NVIDIA AI Podcast by NVIDIA

                                                                                                                                                                NVIDIA AI Podcast

                                                                                                                                                                337 Listeners

                                                                                                                                                                Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                                                Syntax - Tasty Web Development Treats

                                                                                                                                                                985 Listeners

                                                                                                                                                                Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                                                Practical AI

                                                                                                                                                                203 Listeners

                                                                                                                                                                Dwarkesh Podcast by Dwarkesh Patel

                                                                                                                                                                Dwarkesh Podcast

                                                                                                                                                                565 Listeners

                                                                                                                                                                The Data Engineering Show by The Firebolt Data Bros

                                                                                                                                                                The Data Engineering Show

                                                                                                                                                                8 Listeners

                                                                                                                                                                Latent Space: The AI Engineer Podcast by Latent.Space

                                                                                                                                                                Latent Space: The AI Engineer Podcast

                                                                                                                                                                102 Listeners

                                                                                                                                                                This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                                                                                                                                                                This Day in AI Podcast

                                                                                                                                                                222 Listeners

                                                                                                                                                                The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                                                                                                                                                                The AI Daily Brief: Artificial Intelligence News and Analysis

                                                                                                                                                                685 Listeners