Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • Data Migration Strategies For Large Scale Systems
    Summary

    Any software system that survives long enough will require some form of migration or evolution. When that system is responsible for the data layer the process becomes more challenging. Sriram Panyam has been involved in several projects that required migration of large volumes of data in high traffic environments. In this episode he shares some of the valuable lessons that he learned about how to make those projects successful.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst is an end-to-end data lakehouse platform built on Trino, the query engine Apache Iceberg was designed for, with complete support for all table formats including Apache Iceberg, Hive, and Delta Lake. Trusted by teams of all sizes, including Comcast and Doordash. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
    • This episode is supported by Code Comments, an original podcast from Red Hat. As someone who listens to the Data Engineering Podcast, you know that the road from tool selection to production readiness is anything but smooth or straight. In Code Comments, host Jamie Parker, Red Hatter and experienced engineer, shares the journey of technologists from across the industry and their hard-won lessons in implementing new technologies. I listened to the recent episode "Transforming Your Database" and appreciated the valuable advice on how to approach the selection and integration of new databases in applications and the impact on team dynamics. There are 3 seasons of great episodes and new ones landing everywhere you listen to podcasts. Search for "Code Commentst" in your podcast player or go to dataengineeringpodcast.com/codecomments today to subscribe. My thanks to the team at Code Comments for their support.
    • Your host is Tobias Macey and today I'm interviewing Sriram Panyam about his experiences conducting large scale data migrations and the useful strategies that he learned in the process
    • Interview
      • Introduction
      • How did you get involved in the area of data management?
      • Can you start by sharing some of your experiences with data migration projects?
        • As you have gone through successive migration projects, how has that influenced the ways that you think about architecting data systems?
        • How would you categorize the different types and motivations of migrations?
          • How does the motivation for a migration influence the ways that you plan for and execute that work?
          • Can you talk us through one or two specific projects that you have taken part in?
          • Part 1: The Triggers
            • Section 1: Technical Limitations triggering Data Migration
              • Scaling bottlenecks: Performance issues with databases, storage, or network infrastructure
              • Legacy compatibility: Difficulties integrating with modern tools and cloud platforms
              • System upgrades: The need to migrate data during major software changes (e.g., SQL Server version upgrade)
              • Section 2: Types of Migrations for Infrastructure Focus
                • Storage migration: Moving data between systems (HDD to SSD, SAN to NAS, etc.)
                • Data center migration: Physical relocation or consolidation of data centers
                • Virtualization migration: Moving from physical servers to virtual machines (or vice versa)
                • Section 3: Technical Decisions Driving Data Migrations
                  • End-of-life support: Forced migration when older software or hardware is sunsetted
                  • Security and compliance: Adopting new platforms with better security postures
                  • Cost Optimization: Potential savings of cloud vs. on-premise data centers
                  • Part 2: Challenges (and Anxieties)
                    • Section 1: Technical Challenges
                      • Data transformation challenges: Schema changes, complex data mappings
                      • Network bandwidth and latency: Transferring large datasets efficiently
                      • Performance testing and load balancing: Ensuring new systems can handle the workload
                      • Live data consistency: Maintaining data integrity while updates occur in the source system
                      • Minimizing Lag: Techniques to reduce delays in replicating changes to the new system
                      • Change data capture: Identifying and tracking changes to the source system during migration
                      • Section 2: Operational Challenges
                        • Minimizing downtime: Strategies for service continuity during migration
                        • Change management and rollback plans: Dealing with unexpected issues
                        • Technical skills and resources: In-house expertise/data teams/external help
                        • Section 3: Security & Compliance Challenges
                          • Data encryption and protection: Methods for both in-transit and at-rest data
                          • Meeting audit requirements: Documenting data lineage & the chain of custody
                          • Managing access controls: Adjusting identity and role-based access to the new systems
                          • Part 3: Patterns
                            • Section 1: Infrastructure Migration Strategies
                              • Lift and shift: Migrating as-is vs. modernization and re-architecting during the move
                              • Phased vs. big bang approaches: Tradeoffs in risk vs. disruption
                              • Tools and automation: Using specialized software to streamline the process
                              • Dual writes: Managing updates to both old and new systems for a time
                              • Change data capture (CDC) methods: Log-based vs. trigger-based approaches for tracking changes
                              • Data validation & reconciliation: Ensuring consistency between source and target
                              • Section 2: Maintaining Performance and Reliability
                                • Disaster recovery planning: Failover mechanisms for the new environment
                                • Monitoring and alerting: Proactively identifying and addressing issues
                                • Capacity planning and forecasting growth to scale the new infrastructure
                                • Section 3: Data Consistency and Replication
                                  • Replication tools - strategies and specialized tooling
                                  • Data synchronization techniques, eg Pros and cons of different methods (incremental vs. full)
                                  • Testing/Verification Strategies for validating data correctness in a live environment
                                  • Implication of large scale systems/environments
                                  • Comparison of interesting strategies:
                                    • DBLog, Debezium, Databus, Goldengate etc
                                    • What are the most interesting, innovative, or unexpected approaches to data migrations that you have seen or participated in?
                                    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on data migrations?
                                    • When is a migration the wrong choice?
                                    • What are the characteristics or features of data technologies and the overall ecosystem that can reduce the burden of data migration in the future?
                                    • Contact Info
                                      • LinkedIn
                                      • Parting Question
                                        • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                        • Closing Announcements
                                          • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                          • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                          • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                          • Links
                                            • DagKnows
                                            • Google Cloud Dataflow
                                            • Seinfeld Risk Management
                                            • ACL == Access Control List
                                            • LinkedIn Databus - Change Data Capture
                                            • Espresso Storage
                                            • HDFS
                                            • Kafka
                                            • Postgres Replication Slots
                                            • Queueing Theory
                                            • Apache Beam
                                            • Debezium
                                            • Airbyte
                                            • [Fivetran](fivetran.com)
                                            • Designing Data Intensive Applications by Martin Kleppman (affiliate link)
                                            • Vector Databases
                                            • Pinecone
                                            • Weaviate
                                            • LAMP Stack
                                            • Netflix DBLog
                                            • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                              Sponsored By:

                                              • Red Hat Code Comments Podcast: ![Code Comments Podcast Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/A-ygm_NM.jpg)
                                              Putting new technology to use is an exciting prospect. But going from purchase to production isn’t always smooth—even when it’s something everyone is looking forward to. Code Comments covers the bumps, the hiccups, and the setbacks teams face when adjusting to new technology—and the triumphs they pull off once they really get going. Follow Code Comments [anywhere you listen to podcasts](https://link.chtbl.com/codecomments?sid=podcast.dataengineering).
                                            • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                            • This episode is brought to you by Starburst - an end-to-end data lakehouse platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, the query engine Apache Iceberg was designed for, Starburst is an open platform with support for all table formats including Apache Iceberg, Hive, and Delta Lake.
                                              Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. Go to [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)

                                              Support Data Engineering Podcast

                                              1 hr
                                            • Zenlytic Is Building You A Better Coworker With AI Agents
                                              Summary

                                              The purpose of business intelligence systems is to allow anyone in the business to access and decode data to help them make informed decisions. Unfortunately this often turns into an exercise in frustration for everyone involved due to complex workflows and hard-to-understand dashboards. The team at Zenlytic have leaned on the promise of large language models to build an AI agent that lets you converse with your data. In this episode they share their journey through the fast-moving landscape of generative AI and unpack the difference between an AI chatbot and an AI agent.

                                              Announcements
                                              • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                              • This episode is supported by Code Comments, an original podcast from Red Hat. As someone who listens to the Data Engineering Podcast, you know that the road from tool selection to production readiness is anything but smooth or straight. In Code Comments, host Jamie Parker, Red Hatter and experienced engineer, shares the journey of technologists from across the industry and their hard-won lessons in implementing new technologies. I listened to the recent episode "Transforming Your Database" and appreciated the valuable advice on how to approach the selection and integration of new databases in applications and the impact on team dynamics. There are 3 seasons of great episodes and new ones landing everywhere you listen to podcasts. Search for "Code Commentst" in your podcast player or go to dataengineeringpodcast.com/codecomments today to subscribe. My thanks to the team at Code Comments for their support.
                                              • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst is an end-to-end data lakehouse platform built on Trino, the query engine Apache Iceberg was designed for, with complete support for all table formats including Apache Iceberg, Hive, and Delta Lake. Trusted by teams of all sizes, including Comcast and Doordash. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                              • Your host is Tobias Macey and today I'm interviewing Ryan Janssen and Paul Blankley about their experiences building AI powered agents for interacting with your data
                                              • Interview
                                                • Introduction
                                                • How did you get involved in data? In AI?
                                                • Can you describe what Zenlytic is and the role that AI is playing in your platform?
                                                • What have been the key stages in your AI journey?
                                                  • What are some of the dead ends that you ran into along the path to where you are today?
                                                  • What are some of the persistent challenges that you are facing?
                                                  • So tell us more about data agents. Firstly, what are data agents and why do you think they're important?
                                                  • How are data agents different from chatbots?
                                                  • Are data agents harder to build? How do you make them work in production?
                                                  • What other technical architectures have you had to develop to support the use of AI in Zenlytic?
                                                  • How have you approached the work of customer education as you introduce this functionality?
                                                  • What are some of the most interesting or erroneous misconceptions that you have heard about what the AI can and can't do?
                                                  • How have you balanced accuracy/trustworthiness with user experience and flexibility in the conversational AI, given the potential for these models to create erroneous responses?
                                                  • What are the most interesting, innovative, or unexpected ways that you have seen your AI agent used?
                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on building an AI agent for business intelligence?
                                                  • When is an AI agent the wrong choice?
                                                  • What do you have planned for the future of AI in the Zenlytic product?
                                                  • Contact Info
                                                    • Ryan
                                                      • LinkedIn
                                                      • Paul
                                                        • LinkedIn
                                                        • Parting Question
                                                          • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                          • Closing Announcements
                                                            • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                            • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                            • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                            • Links
                                                              • Zenlytic
                                                                • Podcast Episode
                                                                • Attention is all you need
                                                                • Transformers
                                                                • BERT
                                                                • The Bitter Lesson Richard Sutton
                                                                • PID Loops
                                                                • AutoGPT
                                                                • Devin.ai
                                                                • Google Gemini
                                                                • Anthropic Claude
                                                                • OpenAI Code Interpreter
                                                                • Edward Tufte
                                                                • Looker ActionHub
                                                                • OAuth
                                                                • GitHub Copilot
                                                                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                  Sponsored By:

                                                                  • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                  This episode is brought to you by Starburst - an end-to-end data lakehouse platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, the query engine Apache Iceberg was designed for, Starburst is an open platform with support for all table formats including Apache Iceberg, Hive, and Delta Lake.
                                                                  Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. Go to [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                • Red Hat Code Comments Podcast: ![Code Comments Podcast Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/A-ygm_NM.jpg)
                                                                • Putting new technology to use is an exciting prospect. But going from purchase to production isn’t always smooth—even when it’s something everyone is looking forward to. Code Comments covers the bumps, the hiccups, and the setbacks teams face when adjusting to new technology—and the triumphs they pull off once they really get going. Follow Code Comments [anywhere you listen to podcasts](https://link.chtbl.com/codecomments?sid=podcast.dataengineering).

                                                                  Support Data Engineering Podcast

                                                                  55 min
                                                                • Release Management For Data Platform Services And Logic
                                                                  Summary

                                                                  Building a data platform is a substrantial engineering endeavor. Once it is running, the next challenge is figuring out how to address release management for all of the different component parts. The services and systems need to be kept up to date, but so does the code that controls their behavior. In this episode your host Tobias Macey reflects on his current challenges in this area and some of the factors that contribute to the complexity of the problem.

                                                                  Announcements
                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                  • This episode is supported by Code Comments, an original podcast from Red Hat. As someone who listens to the Data Engineering Podcast, you know that the road from tool selection to production readiness is anything but smooth or straight. In Code Comments, host Jamie Parker, Red Hatter and experienced engineer, shares the journey of technologists from across the industry and their hard-won lessons in implementing new technologies. I listened to the recent episode "Transforming Your Database" and appreciated the valuable advice on how to approach the selection and integration of new databases in applications and the impact on team dynamics. There are 3 seasons of great episodes and new ones landing everywhere you listen to podcasts. Search for "Code Commentst" in your podcast player or go to dataengineeringpodcast.com/codecomments today to subscribe. My thanks to the team at Code Comments for their support.
                                                                  • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst is an end-to-end data lakehouse platform built on Trino, the query engine Apache Iceberg was designed for, with complete support for all table formats including Apache Iceberg, Hive, and Delta Lake. Trusted by teams of all sizes, including Comcast and Doordash. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                  • Your host is Tobias Macey and today I want to talk about my experiences managing the QA and release management process of my data platform
                                                                  • Interview
                                                                    • Introduction
                                                                    • As a team, our overall goal is to ensure that the production environment for our data platform is highly stable and reliable. This is the foundational element of establishing and maintaining trust with the consumers of our data. In order to support this effort, we need to ensure that only changes that have been tested and verified are promoted to production.
                                                                    • Our current challenge is one that plagues all data teams. We want to have an environment that mirrors our production environment that is available for testing, but it’s not feasible to maintain a complete duplicate of all of the production data. Compounding that challenge is the fact that each of the components of our data platform interact with data in slightly different ways and need different processes for ensuring that changes are being promoted safely.
                                                                    • Contact Info
                                                                      • LinkedIn
                                                                      • Website
                                                                      • Closing Announcements
                                                                        • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                        • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                        • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
                                                                        • Links
                                                                          • Data Platforms and Leaky Abstractions Episode
                                                                          • Building A Data Platform From Scratch
                                                                          • Airbyte
                                                                            • Podcast Episode
                                                                            • Trino
                                                                            • dbt
                                                                            • Starburst Galaxy
                                                                            • Superset
                                                                            • Dagster
                                                                            • LakeFS
                                                                              • Podcast Episode
                                                                              • Nessie
                                                                                • Podcast Episode
                                                                                • Iceberg
                                                                                • Snowflake
                                                                                • LocalStack
                                                                                • DSL == Domain Specific Language
                                                                                • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                  Sponsored By:

                                                                                  • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                  This episode is brought to you by Starburst - an end-to-end data lakehouse platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, the query engine Apache Iceberg was designed for, Starburst is an open platform with support for all table formats including Apache Iceberg, Hive, and Delta Lake.
                                                                                  Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. Go to [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                • Red Hat Code Comments Podcast: ![Code Comments Podcast Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/A-ygm_NM.jpg)
                                                                                • Putting new technology to use is an exciting prospect. But going from purchase to production isn’t always smooth—even when it’s something everyone is looking forward to. Code Comments covers the bumps, the hiccups, and the setbacks teams face when adjusting to new technology—and the triumphs they pull off once they really get going. Follow Code Comments [anywhere you listen to podcasts](https://link.chtbl.com/codecomments?sid=podcast.dataengineering).

                                                                                  Support Data Engineering Podcast

                                                                                  21 min
                                                                                • Barking Up The Wrong GPTree: Building Better AI With A Cognitive Approach
                                                                                  Summary
                                                                                  Artificial intelligence has dominated the headlines for several months due to the successes of large language models. This has prompted numerous debates about the possibility of, and timeline for, artificial general intelligence (AGI). Peter Voss has dedicated decades of his life to the pursuit of truly intelligent software through the approach of cognitive AI. In this episode he explains his approach to building AI in a more human-like fashion and the emphasis on learning rather than statistical prediction.
                                                                                  Announcements
                                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                  • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                  • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                  • Your host is Tobias Macey and today I'm interviewing Peter Voss about what is involved in making your AI applications more "human"
                                                                                  Interview
                                                                                  • Introduction
                                                                                  • How did you get involved in machine learning?
                                                                                  • Can you start by unpacking the idea of "human-like" AI? 
                                                                                    • How does that contrast with the conception of "AGI"?
                                                                                  • The applications and limitations of GPT/LLM models have been dominating the popular conversation around AI. How do you see that impacting the overrall ecosystem of ML/AI applications and investment?
                                                                                  • The fundamental/foundational challenge of every AI use case is sourcing appropriate data. What are the strategies that you have found useful to acquire, evaluate, and prepare data at an appropriate scale to build high quality models? 
                                                                                  • What are the opportunities and limitations of causal modeling techniques for generalized AI models?
                                                                                  • As AI systems gain more sophistication there is a challenge with establishing and maintaining trust. What are the risks involved in deploying more human-level AI systems and monitoring their reliability?
                                                                                  • What are the practical/architectural methods necessary to build more cognitive AI systems? 
                                                                                    • How would you characterize the ecosystem of tools/frameworks available for creating, evolving, and maintaining these applications?
                                                                                  • What are the most interesting, innovative, or unexpected ways that you have seen cognitive AI applied?
                                                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on desiging/developing cognitive AI systems?
                                                                                  • When is cognitive AI the wrong choice?
                                                                                  • What do you have planned for the future of cognitive AI applications at Aigo?
                                                                                  Contact Info
                                                                                  • LinkedIn
                                                                                  • Website
                                                                                  Parting Question
                                                                                  • From your perspective, what is the biggest barrier to adoption of machine learning today?
                                                                                  Closing Announcements
                                                                                  • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                  • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                  • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                  Links
                                                                                  • Aigo.ai
                                                                                  • Artificial General Intelligence
                                                                                  • Cognitive AI
                                                                                  • Knowledge Graph
                                                                                  • Causal Modeling
                                                                                  • Bayesian Statistics
                                                                                  • Thinking Fast & Slow by Daniel Kahneman (affiliate link)
                                                                                  • Agent-Based Modeling
                                                                                  • Reinforcement Learning
                                                                                  • DARPA 3 Waves of AI presentation
                                                                                  • Why Don't We Have AGI Yet? whitepaper
                                                                                  • Concepts Is All You Need Whitepaper
                                                                                  • Hellen Keller
                                                                                  • Stephen Hawking
                                                                                  The intro and outro music is from Hitman's Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0
                                                                                  55 min
                                                                                • Build Your Second Brain One Piece At A Time
                                                                                  Summary
                                                                                  Generative AI promises to accelerate the productivity of human collaborators. Currently the primary way of working with these tools is through a conversational prompt, which is often cumbersome and unwieldy. In order to simplify the integration of AI capabilities into developer workflows Tsavo Knott helped create Pieces, a powerful collection of tools that complements the tools that developers already use. In this episode he explains the data collection and preparation process, the collection of model types and sizes that work together to power the experience, and how to incorporate it into your workflow to act as a second brain.


                                                                                  Announcements
                                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                  • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                  • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                  • Your host is Tobias Macey and today I'm interviewing Tsavo Knott about Pieces, a personal AI toolkit to improve the efficiency of developers
                                                                                  Interview
                                                                                  • Introduction
                                                                                  • How did you get involved in machine learning?
                                                                                  • Can you describe what Pieces is and the story behind it?
                                                                                  • The past few months have seen an endless series of personalized AI tools launched. What are the features and focus of Pieces that might encourage someone to use it over the alternatives?
                                                                                  • model selections
                                                                                  • architecture of Pieces application
                                                                                  • local vs. hybrid vs. online models
                                                                                  • model update/delivery process
                                                                                  • data preparation/serving for models in context of Pieces app
                                                                                  • application of AI to developer workflows
                                                                                  • types of workflows that people are building with pieces
                                                                                  • What are the most interesting, innovative, or unexpected ways that you have seen Pieces used?
                                                                                  • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Pieces?
                                                                                  • When is Pieces the wrong choice?
                                                                                  • What do you have planned for the future of Pieces?
                                                                                  Contact Info
                                                                                  • LinkedIn
                                                                                  Parting Question
                                                                                  • From your perspective, what is the biggest barrier to adoption of machine learning today?
                                                                                  Closing Announcements
                                                                                  • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                  • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                  • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                  Links
                                                                                  • Pieces
                                                                                  • NPU == Neural Processing Unit
                                                                                  • Tensor Chip
                                                                                  • LoRA == Low Rank Adaptation
                                                                                  • Generative Adversarial Networks
                                                                                  • Mistral
                                                                                  • Emacs
                                                                                  • Vim
                                                                                  • NeoVim
                                                                                  • Dart
                                                                                  • Flutter
                                                                                  • Typescript
                                                                                  • Lua
                                                                                  • Retrieval Augmented Generation
                                                                                  • ONNX
                                                                                  • LSTM == Long Short-Term Memory
                                                                                  • LLama 2
                                                                                  • GitHub Copilot
                                                                                  • Tabnine
                                                                                    • Podcast Episode
                                                                                  The intro and outro music is from Hitman's Lovesong feat. Paola Graziano by The Freak Fandango Orchestra/CC BY-SA 3.0
                                                                                  51 min
                                                                                • Making Email Better With AI At Shortwave
                                                                                  Summary

                                                                                  Generative AI has rapidly transformed everything in the technology sector. When Andrew Lee started work on Shortwave he was focused on making email more productive. When AI started gaining adoption he realized that he had even more potential for a transformative experience. In this episode he shares the technical challenges that he and his team have overcome in integrating AI into their product, as well as the benefits and features that it provides to their customers.

                                                                                  Announcements
                                                                                  • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                  • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                  • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                  • Your host is Tobias Macey and today I'm interviewing Andrew Lee about his work on Shortwave, an AI powered email client
                                                                                  • Interview
                                                                                    • Introduction
                                                                                    • How did you get involved in the area of data management?
                                                                                    • Can you describe what Shortwave is and the story behind it?
                                                                                      • What is the core problem that you are addressing with Shortwave?
                                                                                      • Email has been a central part of communication and business productivity for decades now. What are the overall themes that continue to be problematic?
                                                                                      • What are the strengths that email maintains as a protocol and ecosystem?
                                                                                      • From a product perspective, what are the data challenges that are posed by email?
                                                                                      • Can you describe how you have architected the Shortwave platform?
                                                                                        • How have the design and goals of the product changed since you started it?
                                                                                        • What are the ways that the advent and evolution of language models have influenced your product roadmap?
                                                                                        • How do you manage the personalization of the AI functionality in your system for each user/team?
                                                                                        • For users and teams who are using Shortwave, how does it change their workflow and communication patterns?
                                                                                        • Can you describe how I would use Shortwave for managing the workflow of evaluating, planning, and promoting my podcast episodes?
                                                                                        • What are the most interesting, innovative, or unexpected ways that you have seen Shortwave used?
                                                                                        • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Shortwave?
                                                                                        • When is Shortwave the wrong choice?
                                                                                        • What do you have planned for the future of Shortwave?
                                                                                        • Contact Info
                                                                                          • LinkedIn
                                                                                          • Blog
                                                                                          • Parting Question
                                                                                            • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                            • Closing Announcements
                                                                                              • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                              • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                              • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                              • Links
                                                                                                • Shortwave
                                                                                                • Firebase
                                                                                                • Google Inbox
                                                                                                • Hey
                                                                                                  • Ezra Klein Hey Article
                                                                                                  • Superhuman
                                                                                                  • Pinecone
                                                                                                    • Podcast Episode
                                                                                                    • Elastic
                                                                                                    • Hybrid Search
                                                                                                    • Semantic Search
                                                                                                    • Mistral
                                                                                                    • GPT 3.5
                                                                                                    • IMAP
                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                      Sponsored By:

                                                                                                      • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                                      This episode is brought to you by Starburst - a data lake analytics platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, Starburst runs petabyte-scale SQL analytics fast at a fraction of the cost of traditional methods, helping you meet all your data needs ranging from AI/ML workloads to data applications to complete analytics.
                                                                                                      Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                                    • Dagster: ![Dagster Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/jz4xfquZ.png)
                                                                                                    • Data teams are tasked with helping organizations deliver on the premise of data, and with ML and AI maturing rapidly, expectations have never been this high. However data engineers are challenged by both technical complexity and organizational complexity, with heterogeneous technologies to adopt, multiple data disciplines converging, legacy systems to support, and costs to manage.
                                                                                                      Dagster is an open-source orchestration solution that helps data teams reign in this complexity and build data platforms that provide unparalleled observability, and testability, all while fostering collaboration across the enterprise. With enterprise-grade hosting on Dagster Cloud, you gain even more capabilities, adding cost management, security, and CI support to further boost your teams' productivity. Go to [dagster.io](https://dagster.io/lp/dagster-cloud-trial?source=data-eng-podcast) today to get your first 30 days free!

                                                                                                      Support Data Engineering Podcast

                                                                                                      54 min
                                                                                                    • Designing A Non-Relational Database Engine
                                                                                                      Summary

                                                                                                      Databases come in a variety of formats for different use cases. The default association with the term "database" is relational engines, but non-relational engines are also used quite widely. In this episode Oren Eini, CEO and creator of RavenDB, explores the nuances of relational vs. non-relational engines, and the strategies for designing a non-relational database.

                                                                                                      Announcements
                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                      • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting dataengineeringpodcast.com/datafold.
                                                                                                      • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                                      • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                                      • Your host is Tobias Macey and today I'm interviewing Oren Eini about the work of designing and building a NoSQL database engine
                                                                                                      • Interview
                                                                                                        • Introduction
                                                                                                        • How did you get involved in the area of data management?
                                                                                                        • Can you describe what constitutes a NoSQL database?
                                                                                                          • How have the requirements and applications of NoSQL engines changed since they first became popular ~15 years ago?
                                                                                                          • What are the factors that convince teams to use a NoSQL vs. SQL database?
                                                                                                            • NoSQL is a generalized term that encompasses a number of different data models. How does the underlying representation (e.g. document, K/V, graph) change that calculus?
                                                                                                            • How have the evolution in data formats (e.g. N-dimensional vectors, point clouds, etc.) changed the landscape for NoSQL engines?
                                                                                                            • When designing and building a database, what are the initial set of questions that need to be answered?
                                                                                                              • How many "core capabilities" can you reasonably design around before they conflict with each other?
                                                                                                              • How have you approached the evolution of RavenDB as you add new capabilities and mature the project?
                                                                                                                • What are some of the early decisions that had to be unwound to enable new capabilities?
                                                                                                                • If you were to start from scratch today, what database would you build?
                                                                                                                • What are the most interesting, innovative, or unexpected ways that you have seen RavenDB/NoSQL databases used?
                                                                                                                • What are the most interesting, unexpected, or challenging lessons that you have learned while working on RavenDB?
                                                                                                                • When is a NoSQL database/RavenDB the wrong choice?
                                                                                                                • What do you have planned for the future of RavenDB?
                                                                                                                • Contact Info
                                                                                                                  • Blog
                                                                                                                  • LinkedIn
                                                                                                                  • Parting Question
                                                                                                                    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                    • Closing Announcements
                                                                                                                      • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                      • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                      • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                      • Links
                                                                                                                        • RavenDB
                                                                                                                        • RSS
                                                                                                                        • Object Relational Mapper (ORM)
                                                                                                                        • Relational Database
                                                                                                                        • NoSQL
                                                                                                                        • CouchDB
                                                                                                                        • Navigational Database
                                                                                                                        • MongoDB
                                                                                                                        • Redis
                                                                                                                        • Neo4J
                                                                                                                        • Cassandra
                                                                                                                        • Column-Family
                                                                                                                        • SQLite
                                                                                                                        • LevelDB
                                                                                                                        • Firebird DB
                                                                                                                        • fsync
                                                                                                                        • Esent DB?
                                                                                                                        • KNN == K-Nearest Neighbors
                                                                                                                        • RocksDB
                                                                                                                        • C# Language
                                                                                                                        • ASP.NET
                                                                                                                        • QUIC
                                                                                                                        • Dynamo Paper
                                                                                                                        • Database Internals book (affiliate link)
                                                                                                                        • Designing Data Intensive Applications book (affiliate link)
                                                                                                                        • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                          Sponsored By:

                                                                                                                          • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                                                          This episode is brought to you by Starburst - a data lake analytics platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, Starburst runs petabyte-scale SQL analytics fast at a fraction of the cost of traditional methods, helping you meet all your data needs ranging from AI/ML workloads to data applications to complete analytics.
                                                                                                                          Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                                                        • Datafold: ![Datafold](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zm6x2tFu.png)
                                                                                                                        • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting https://get.datafold.com/replication-de-podcast.
                                                                                                                        • Dagster: ![Dagster Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/jz4xfquZ.png)
                                                                                                                        • Data teams are tasked with helping organizations deliver on the premise of data, and with ML and AI maturing rapidly, expectations have never been this high. However data engineers are challenged by both technical complexity and organizational complexity, with heterogeneous technologies to adopt, multiple data disciplines converging, legacy systems to support, and costs to manage.
                                                                                                                          Dagster is an open-source orchestration solution that helps data teams reign in this complexity and build data platforms that provide unparalleled observability, and testability, all while fostering collaboration across the enterprise. With enterprise-grade hosting on Dagster Cloud, you gain even more capabilities, adding cost management, security, and CI support to further boost your teams' productivity. Go to [dagster.io](https://dagster.io/lp/dagster-cloud-trial?source=data-eng-podcast) today to get your first 30 days free!

                                                                                                                          Support Data Engineering Podcast

                                                                                                                          1 hr 17 min
                                                                                                                        • Establish A Single Source Of Truth For Your Data Consumers With A Semantic Layer
                                                                                                                          Summary

                                                                                                                          Maintaining a single source of truth for your data is the biggest challenge in data engineering. Different roles and tasks in the business need their own ways to access and analyze the data in the organization. In order to enable this use case, while maintaining a single point of access, the semantic layer has evolved as a technological solution to the problem. In this episode Artyom Keydunov, creator of Cube, discusses the evolution and applications of the semantic layer as a component of your data platform, and how Cube provides speed and cost optimization for your data consumers.

                                                                                                                          Announcements
                                                                                                                          • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                          • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting dataengineeringpodcast.com/datafold.
                                                                                                                          • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                                                          • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                                                          • Your host is Tobias Macey and today I'm interviewing Artyom Keydunov about the role of the semantic layer in your data platform
                                                                                                                          • Interview
                                                                                                                            • Introduction
                                                                                                                            • How did you get involved in the area of data management?
                                                                                                                            • Can you start by outlining the technical elements of what it means to have a "semantic layer"?
                                                                                                                            • In the past couple of years there was a rapid hype cycle around the "metrics layer" and "headless BI", which has largely faded. Can you give your assessment of the current state of the industry around the adoption/implementation of these concepts?
                                                                                                                            • What are the benefits of having a discrete service that offers the business metrics/semantic mappings as opposed to implementing those concepts as part of a more general system? (e.g. dbt, BI, warehouse marts, etc.)
                                                                                                                              • At what point does it become necessary/beneficial for a team to adopt such a service?
                                                                                                                              • What are the challenges involved in retrofitting a semantic layer into a production data system?
                                                                                                                              • evolution of requirements/usage patterns
                                                                                                                              • technical complexities/performance and cost optimization
                                                                                                                              • What are the most interesting, innovative, or unexpected ways that you have seen Cube used?
                                                                                                                              • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Cube?
                                                                                                                              • When is Cube/a semantic layer the wrong choice?
                                                                                                                              • What do you have planned for the future of Cube?
                                                                                                                              • Contact Info
                                                                                                                                • LinkedIn
                                                                                                                                • keydunov on GitHub
                                                                                                                                • Parting Question
                                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                  • Closing Announcements
                                                                                                                                    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                    • Links
                                                                                                                                      • Cube
                                                                                                                                      • Semantic Layer
                                                                                                                                      • Business Objects
                                                                                                                                      • Tableau
                                                                                                                                      • Looker
                                                                                                                                        • Podcast Episode
                                                                                                                                        • Mode
                                                                                                                                        • Thoughtspot
                                                                                                                                        • LightDash
                                                                                                                                          • Podcast Episode
                                                                                                                                          • Embedded Analytics
                                                                                                                                          • Dimensional Modeling
                                                                                                                                          • Clickhouse
                                                                                                                                            • Podcast Episode
                                                                                                                                            • Druid
                                                                                                                                            • BigQuery
                                                                                                                                            • Starburst
                                                                                                                                            • Pinot
                                                                                                                                            • Snowflake
                                                                                                                                              • Podcast Episode
                                                                                                                                              • Arrow Datafusion
                                                                                                                                              • Metabase
                                                                                                                                                • Podcast Episode
                                                                                                                                                • Superset
                                                                                                                                                • Alation
                                                                                                                                                • Collibra
                                                                                                                                                  • Podcast Episode
                                                                                                                                                  • Atlan
                                                                                                                                                    • Podcast Episode
                                                                                                                                                    • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                      Sponsored By:

                                                                                                                                                      • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                                                                                      This episode is brought to you by Starburst - a data lake analytics platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, Starburst runs petabyte-scale SQL analytics fast at a fraction of the cost of traditional methods, helping you meet all your data needs ranging from AI/ML workloads to data applications to complete analytics.
                                                                                                                                                      Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                                                                                    • Datafold: ![Datafold](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zm6x2tFu.png)
                                                                                                                                                    • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting https://get.datafold.com/replication-de-podcast.
                                                                                                                                                    • Dagster: ![Dagster Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/jz4xfquZ.png)
                                                                                                                                                    • Data teams are tasked with helping organizations deliver on the premise of data, and with ML and AI maturing rapidly, expectations have never been this high. However data engineers are challenged by both technical complexity and organizational complexity, with heterogeneous technologies to adopt, multiple data disciplines converging, legacy systems to support, and costs to manage.
                                                                                                                                                      Dagster is an open-source orchestration solution that helps data teams reign in this complexity and build data platforms that provide unparalleled observability, and testability, all while fostering collaboration across the enterprise. With enterprise-grade hosting on Dagster Cloud, you gain even more capabilities, adding cost management, security, and CI support to further boost your teams' productivity. Go to [dagster.io](https://dagster.io/lp/dagster-cloud-trial?source=data-eng-podcast) today to get your first 30 days free!

                                                                                                                                                      Support Data Engineering Podcast

                                                                                                                                                      57 min
                                                                                                                                                    • Adding Anomaly Detection And Observability To Your dbt Projects Is Elementary
                                                                                                                                                      Summary

                                                                                                                                                      Working with data is a complicated process, with numerous chances for something to go wrong. Identifying and accounting for those errors is a critical piece of building trust in the organization that your data is accurate and up to date. While there are numerous products available to provide that visibility, they all have different technologies and workflows that they focus on. To bring observability to dbt projects the team at Elementary embedded themselves into the workflow. In this episode Maayan Salom explores the approach that she has taken to bring observability, enhanced testing capabilities, and anomaly detection into every step of the dbt developer experience.

                                                                                                                                                      Announcements
                                                                                                                                                      • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                      • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                                                                                      • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                                                                                      • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting dataengineeringpodcast.com/datafold.
                                                                                                                                                      • Your host is Tobias Macey and today I'm interviewing Maayan Salom about how to incorporate observability into a dbt-oriented workflow and how Elementary can help
                                                                                                                                                      • Interview
                                                                                                                                                        • Introduction
                                                                                                                                                        • How did you get involved in the area of data management?
                                                                                                                                                        • Can you start by outlining what elements of observability are most relevant for dbt projects?
                                                                                                                                                        • What are some of the common ad-hoc/DIY methods that teams develop to acquire those insights?
                                                                                                                                                          • What are the challenges/shortcomings associated with those approaches?
                                                                                                                                                          • Over the past ~3 years there were numerous data observability systems/products created. What are some of the ways that the specifics of dbt workflows are not covered by those generalized tools?
                                                                                                                                                            • What are the insights that can be more easily generated by embedding into the dbt toolchain and development cycle?
                                                                                                                                                            • Can you describe what Elementary is and how it is designed to enhance the development and maintenance work in dbt projects?
                                                                                                                                                            • How is Elementary designed/implemented?
                                                                                                                                                              • How have the scope and goals of the project changed since you started working on it?
                                                                                                                                                              • What are the engineering challenges/frustrations that you have dealt with in the creation and evolution of Elementary?
                                                                                                                                                              • Can you talk us through the setup and workflow for teams adopting Elementary in their dbt projects?
                                                                                                                                                              • How does the incorporation of Elementary change the development habits of the teams who are using it?
                                                                                                                                                              • What are the most interesting, innovative, or unexpected ways that you have seen Elementary used?
                                                                                                                                                              • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Elementary?
                                                                                                                                                              • When is Elementary the wrong choice?
                                                                                                                                                              • What do you have planned for the future of Elementary?
                                                                                                                                                              • Contact Info
                                                                                                                                                                • LinkedIn
                                                                                                                                                                • Parting Question
                                                                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                  • Closing Announcements
                                                                                                                                                                    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                    • Links
                                                                                                                                                                      • Elementary
                                                                                                                                                                      • Data Observability
                                                                                                                                                                      • dbt
                                                                                                                                                                      • Datadog
                                                                                                                                                                      • pre-commit
                                                                                                                                                                      • dbt packages
                                                                                                                                                                      • SQLMesh
                                                                                                                                                                      • Malloy
                                                                                                                                                                      • SDF
                                                                                                                                                                      • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                        Sponsored By:

                                                                                                                                                                        • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                                                                                                        This episode is brought to you by Starburst - a data lake analytics platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, Starburst runs petabyte-scale SQL analytics fast at a fraction of the cost of traditional methods, helping you meet all your data needs ranging from AI/ML workloads to data applications to complete analytics.
                                                                                                                                                                        Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                                                                                                      • Datafold: ![Datafold](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/zm6x2tFu.png)
                                                                                                                                                                      • This episode is brought to you by Datafold – a testing automation platform for data engineers that prevents data quality issues from entering every part of your data workflow, from migration to dbt deployment. Datafold has recently launched data replication testing, providing ongoing validation for source-to-target replication. Leverage Datafold's fast cross-database data diffing and Monitoring to test your replication pipelines automatically and continuously. Validate consistency between source and target at any scale, and receive alerts about any discrepancies. Learn more about Datafold by visiting https://get.datafold.com/replication-de-podcast.
                                                                                                                                                                      • Dagster: ![Dagster Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/jz4xfquZ.png)
                                                                                                                                                                      • Data teams are tasked with helping organizations deliver on the premise of data, and with ML and AI maturing rapidly, expectations have never been this high. However data engineers are challenged by both technical complexity and organizational complexity, with heterogeneous technologies to adopt, multiple data disciplines converging, legacy systems to support, and costs to manage.
                                                                                                                                                                        Dagster is an open-source orchestration solution that helps data teams reign in this complexity and build data platforms that provide unparalleled observability, and testability, all while fostering collaboration across the enterprise. With enterprise-grade hosting on Dagster Cloud, you gain even more capabilities, adding cost management, security, and CI support to further boost your teams' productivity. Go to [dagster.io](https://dagster.io/lp/dagster-cloud-trial?source=data-eng-podcast) today to get your first 30 days free!

                                                                                                                                                                        Support Data Engineering Podcast

                                                                                                                                                                        51 min
                                                                                                                                                                      • Ship Smarter Not Harder With Declarative And Collaborative Data Orchestration On Dagster+
                                                                                                                                                                        Summary

                                                                                                                                                                        A core differentiator of Dagster in the ecosystem of data orchestration is their focus on software defined assets as a means of building declarative workflows. With their launch of Dagster+ as the redesigned commercial companion to the open source project they are investing in that capability with a suite of new features. In this episode Pete Hunt, CEO of Dagster labs, outlines these new capabilities, how they reduce the burden on data teams, and the increased collaboration that they enable across teams and business units.

                                                                                                                                                                        Announcements
                                                                                                                                                                        • Hello and welcome to the Data Engineering Podcast, the show about modern data management
                                                                                                                                                                        • Dagster offers a new approach to building and running data platforms and data pipelines. It is an open-source, cloud-native orchestrator for the whole development lifecycle, with integrated lineage and observability, a declarative programming model, and best-in-class testability. Your team can get up and running in minutes thanks to Dagster Cloud, an enterprise-class hosted solution that offers serverless and hybrid deployments, enhanced security, and on-demand ephemeral test deployments. Go to dataengineeringpodcast.com/dagster today to get started. Your first 30 days are free!
                                                                                                                                                                        • Data lakes are notoriously complex. For data engineers who battle to build and scale high quality data workflows on the data lake, Starburst powers petabyte-scale SQL analytics fast, at a fraction of the cost of traditional methods, so that you can meet all your data needs ranging from AI to data applications to complete analytics. Trusted by teams of all sizes, including Comcast and Doordash, Starburst is a data lake analytics platform that delivers the adaptability and flexibility a lakehouse ecosystem promises. And Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Go to dataengineeringpodcast.com/starburst and get $500 in credits to try Starburst Galaxy today, the easiest and fastest way to get started using Trino.
                                                                                                                                                                        • Your host is Tobias Macey and today I'm interviewing Pete Hunt about how the launch of Dagster+ will level up your data platform and orchestrate across language platforms
                                                                                                                                                                        • Interview
                                                                                                                                                                          • Introduction
                                                                                                                                                                          • How did you get involved in the area of data management?
                                                                                                                                                                          • Can you describe what the focus of Dagster+ is and the story behind it?
                                                                                                                                                                            • What problems are you trying to solve with Dagster+?
                                                                                                                                                                            • What are the notable enhancements beyond the Dagster Core project that this updated platform provides?
                                                                                                                                                                            • How is it different from the current Dagster Cloud product?
                                                                                                                                                                            • In the launch announcement you tease new capabilities that would be great to explore in turns:
                                                                                                                                                                              • Make data a team sport, enabling data teams across the organization
                                                                                                                                                                              • Deliver reliable, high quality data the organization can trust
                                                                                                                                                                              • Observe and manage data platform costs
                                                                                                                                                                              • Master the heterogeneous collection of technologies—both traditional and Modern Data Stack
                                                                                                                                                                              • What are the business/product goals that you are focused on improving with the launch of Dagster+
                                                                                                                                                                              • What are the most interesting, innovative, or unexpected ways that you have seen Dagster used?
                                                                                                                                                                              • What are the most interesting, unexpected, or challenging lessons that you have learned while working on the design and launch of Dagster+?
                                                                                                                                                                              • When is Dagster+ the wrong choice?
                                                                                                                                                                              • What do you have planned for the future of Dagster/Dagster Cloud/Dagster+?
                                                                                                                                                                              • Contact Info
                                                                                                                                                                                • Twitter
                                                                                                                                                                                • LinkedIn
                                                                                                                                                                                • Parting Question
                                                                                                                                                                                  • From your perspective, what is the biggest gap in the tooling or technology for data management today?
                                                                                                                                                                                  • Closing Announcements
                                                                                                                                                                                    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The Machine Learning Podcast helps you go from idea to production with machine learning.
                                                                                                                                                                                    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
                                                                                                                                                                                    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected]) with your story.
                                                                                                                                                                                    • Links
                                                                                                                                                                                      • Dagster
                                                                                                                                                                                        • Podcast Episode
                                                                                                                                                                                        • Dagster+ Launch Event
                                                                                                                                                                                        • Hadoop
                                                                                                                                                                                        • MapReduce
                                                                                                                                                                                        • Pydantic
                                                                                                                                                                                        • Software Defined Assets
                                                                                                                                                                                        • Dagster Insights
                                                                                                                                                                                        • Dagster Pipes
                                                                                                                                                                                        • Conway's Law
                                                                                                                                                                                        • Data Mesh
                                                                                                                                                                                        • Dagster Code Locations
                                                                                                                                                                                        • Dagster Asset Checks
                                                                                                                                                                                        • Dave & Buster's
                                                                                                                                                                                        • SQLMesh
                                                                                                                                                                                          • Podcast Episode
                                                                                                                                                                                          • SDF
                                                                                                                                                                                          • Malloy
                                                                                                                                                                                          • The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA

                                                                                                                                                                                            Sponsored By:

                                                                                                                                                                                            • Starburst: ![Starburst Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/UpvN7wDT.png)
                                                                                                                                                                                            This episode is brought to you by Starburst - a data lake analytics platform for data engineers who are battling to build and scale high quality data pipelines on the data lake. Powered by Trino, Starburst runs petabyte-scale SQL analytics fast at a fraction of the cost of traditional methods, helping you meet all your data needs ranging from AI/ML workloads to data applications to complete analytics.
                                                                                                                                                                                            Trusted by the teams at Comcast and Doordash, Starburst delivers the adaptability and flexibility a lakehouse ecosystem promises, while providing a single point of access for your data and all your data governance allowing you to discover, transform, govern, and secure all in one place. Starburst does all of this on an open architecture with first-class support for Apache Iceberg, Delta Lake and Hudi, so you always maintain ownership of your data. Want to see Starburst in action? Try Starburst Galaxy today, the easiest and fastest way to get started using Trino, and get $500 of credits free. [dataengineeringpodcast.com/starburst](https://www.dataengineeringpodcast.com/starburst)
                                                                                                                                                                                          • Dagster: ![Dagster Logo](https://files.fireside.fm/file/fireside-uploads/images/c/c6161a3f-a67b-48ef-b087-52f1f1573292/jz4xfquZ.png)
                                                                                                                                                                                          • Data teams are tasked with helping organizations deliver on the premise of data, and with ML and AI maturing rapidly, expectations have never been this high. However data engineers are challenged by both technical complexity and organizational complexity, with heterogeneous technologies to adopt, multiple data disciplines converging, legacy systems to support, and costs to manage.
                                                                                                                                                                                            Dagster is an open-source orchestration solution that helps data teams reign in this complexity and build data platforms that provide unparalleled observability, and testability, all while fostering collaboration across the enterprise. With enterprise-grade hosting on Dagster Cloud, you gain even more capabilities, adding cost management, security, and CI support to further boost your teams' productivity. Go to [dagster.io](https://dagster.io/lp/dagster-cloud-trial?source=data-eng-podcast) today to get your first 30 days free!

                                                                                                                                                                                            Support Data Engineering Podcast

                                                                                                                                                                                            56 min

                                                                                                                                                                                          About Data Engineering Podcast

                                                                                                                                                                                          From the publisher's feed

                                                                                                                                                                                          This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

                                                                                                                                                                                          More shows like Data Engineering Podcast

                                                                                                                                                                                          This Week in Startups by Jason Calacanis

                                                                                                                                                                                          This Week in Startups

                                                                                                                                                                                          1,290 Listeners

                                                                                                                                                                                          The Changelog: Software Development, Open Source by Changelog Media

                                                                                                                                                                                          The Changelog: Software Development, Open Source

                                                                                                                                                                                          286 Listeners

                                                                                                                                                                                          The a16z Show by Andreessen Horowitz

                                                                                                                                                                                          The a16z Show

                                                                                                                                                                                          1,087 Listeners

                                                                                                                                                                                          Software Engineering Daily by Software Engineering Daily

                                                                                                                                                                                          Software Engineering Daily

                                                                                                                                                                                          623 Listeners

                                                                                                                                                                                          Risky Business by Risky Business Media

                                                                                                                                                                                          Risky Business

                                                                                                                                                                                          375 Listeners

                                                                                                                                                                                          Talk Python To Me by Michael Kennedy

                                                                                                                                                                                          Talk Python To Me

                                                                                                                                                                                          582 Listeners

                                                                                                                                                                                          Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

                                                                                                                                                                                          Super Data Science: ML & AI Podcast with Jon Krohn

                                                                                                                                                                                          305 Listeners

                                                                                                                                                                                          NVIDIA AI Podcast by NVIDIA

                                                                                                                                                                                          NVIDIA AI Podcast

                                                                                                                                                                                          338 Listeners

                                                                                                                                                                                          Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

                                                                                                                                                                                          Syntax - Tasty Web Development Treats

                                                                                                                                                                                          985 Listeners

                                                                                                                                                                                          Practical AI by Daniel Whitenack and Chris Benson

                                                                                                                                                                                          Practical AI

                                                                                                                                                                                          203 Listeners

                                                                                                                                                                                          Dwarkesh Podcast by Dwarkesh Patel

                                                                                                                                                                                          Dwarkesh Podcast

                                                                                                                                                                                          565 Listeners

                                                                                                                                                                                          The Data Engineering Show by The Firebolt Data Bros

                                                                                                                                                                                          The Data Engineering Show

                                                                                                                                                                                          8 Listeners

                                                                                                                                                                                          Latent Space: The AI Engineer Podcast by Latent.Space

                                                                                                                                                                                          Latent Space: The AI Engineer Podcast

                                                                                                                                                                                          102 Listeners

                                                                                                                                                                                          This Day in AI Podcast by Michael Sharkey, Chris Sharkey

                                                                                                                                                                                          This Day in AI Podcast

                                                                                                                                                                                          222 Listeners

                                                                                                                                                                                          The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

                                                                                                                                                                                          The AI Daily Brief: Artificial Intelligence News and Analysis

                                                                                                                                                                                          684 Listeners