Data Engineering Podcast

Data Engineering Podcast

By Tobias MaceyTechnologyEducation
Download on the App Store
  • Favorites

    137

    Followers

  • Typical duration

    56 min

    per episode

Based on Podcast App listening data

Data Engineering Podcast episodes

  • High Performance And Low Overhead Graphs With KuzuDB
    Summary
    In this episode of the Data Engineering Podcast Prashanth Rao, an AI engineer at KuzuDB, talks about their embeddable graph database. Prashanth explains how KuzuDB addresses performance shortcomings in existing solutions through columnar storage and novel join algorithms. He discusses the usability and scalability of KuzuDB, emphasizing its open-source nature and potential for various graph applications. The conversation explores the growing interest in graph databases due to their AI and data engineering applications, and Prashanth highlights KuzuDB's potential in edge computing, ephemeral workloads, and integration with other formats like Iceberg and Parquet.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Prashanth Rao about KuzuDB, an embeddable graph database
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what KuzuDB is and the story behind it?
    • What are the core use cases that Kuzu is focused on addressing?
      • What is explicitly out of scope?
    • Graph engines have been available and in use for a long time, but generally for more niche use cases. How would you characterize the current state of the graph data ecosystem?
    • You note scalability as a feature of Kuzu, which is a phrase with many potential interpretations. Typically horizontal scaling of graphs has been complicated, in what sense does Kuzu make that claim?
    • Can you describe some of the typical architecture and integration patterns of Kuzu?
      • What are some of the more interesting or esoteric means of architecting with Kuzu?
    • For cases where Kuzu is rendering a graph across an external data repository (e.g. Iceberg, etc.), what are the patterns for balancing data freshness with network/compute efficiency? (e.g. read and create every time or persist the Kuzu state)
    • Can you describe the internal architecture of Kuzu and key design factors?
      • What are the benefits and tradeoffs of using a columnar store with adjacency lists vs. a more graph-native storage format?
    • What are the most interesting, innovative, or unexpected ways that you have seen Kuzu used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Kuzu?
    • When is Kuzu the wrong choice?
    • What do you have planned for the future of Kuzu?
    Contact Info
    • Website
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Links
    • KuzuDB
    • BERT
    • Transformer Architecture
    • DuckDB
      • Podcast Episode
    • MonetDB
    • Umbra DB
    • sqlite
    • Cypher Query Language
    • Property Graph
    • Neo4J
    • GraphRAG
    • Context Engineering
    • Write-Ahead Log
    • Bauplan
    • Iceberg
    • DuckLake
    • Lance
    • LanceDB
    • Arrow
    • Polars
    • Arrow DataFusion
    • GQL
    • ClickHouse
    • Adjacency List
    • Why Graph Databases Need New Join Algorithms
    • KuzuDB WASM
    • RAG == Retrieval Augmented Generation
    • NetworkX
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 2 min
  • Bridging Data and Decision-Making: AI's Role in Modern Analytics
    Summary
    In this episode of the Data Engineering Podcast Lucas Thelosen and Drew Gilson from Gravity talk about their development of Orion, an autonomous data analyst that bridges the gap between data availability and business decision-making. Lucas and Drew share their backgrounds in data analytics and how their experiences have shaped their approach to leveraging AI for data analysis, emphasizing the potential of AI to democratize data insights and make sophisticated analysis accessible to companies of all sizes. They discuss the technical aspects of Orion, a multi-agent system designed to automate data analysis and provide actionable insights, highlighting the importance of integrating AI into existing workflows with accuracy and trustworthiness in mind. The conversation also explores how AI can free data analysts from routine tasks, enabling them to focus on strategic decision-making and stakeholder management, as they discuss the future of AI in data analytics and its transformative impact on businesses.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Lucas Thelosen and Drew Gilson about the engineering and impact of building an autonomous data analyst
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Orion is and the story behind it?
      • How do you envision the role of an agentic analyst in an organizational context?
    • There have been several attempts at building LLM-powered data analysis, many of which are essentially a text-to-SQL interface. How have the capabilities and architectural patterns grown in the past ~2 years to enable a more capable system?
    • One of the key success factors for a data analyst is their ability to translate business questions into technical representations. How can an autonomous AI-powered system understand the complex nuance of the business to build effective analyses?
    • Many agentic approaches to analytics require a substantial investment in data architecture, documentation, and semantic models to be effective. What are the gradations of effectiveness for autonomous analytics for companies who are at different points on their journey to technical maturity?
    • Beyond raw capability, there is also a significant need to invest in user experience design for an agentic analyst to be useful. What are the key interaction patterns that you have found to be helpful as you have developed your system?
    • How does the introduction of a system like Orion shift the workload for data teams?
    • Can you describe the overall system design and technical architecture of Orion?
      • How has that changed as you gained further experience and understanding of the problem space?
    • What are the most interesting, innovative, or unexpected ways that you have seen Orion used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Orion?
    • When is Orion/agentic analytics the wrong choice?
    • What do you have planned for the future of Orion?
    Contact Info
    • Lucas
      • LinkedIn
    • Drew
      • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Orion
    • Looker
    • Gravity
    • VBA == Visual Basic for Applications
    • Text-To-SQL
    • One-shot
    • LookML
    • Data Grain
    • LLM As A Judge
    • Google Large Time Series Model
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 11 min
  • From Bits to Tables: The Evolution of S3 Storage
    Summary
    In this episode of the Data Engineering Podcast Andy Warfield talks about the innovative functionalities of S3 Tables and Vectors and their integration into modern data stacks. Andy shares his journey through the tech industry and his role at Amazon, where he collaborates to enhance storage capabilities, discussing the evolution of S3 from a simple storage solution to a sophisticated system supporting advanced data types like tables and vectors crucial for analytics and AI-driven applications. He explains the motivations behind introducing S3 Tables and Vectors, highlighting their role in simplifying data management and enhancing performance for complex workloads, and shares insights into the technical challenges and design considerations involved in developing these features. The conversation explores potential applications of S3 Tables and Vectors in fields like AI, genomics, and media, and discusses future directions for S3's development to further support data-driven innovation.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Tired of data migrations that drag on for months or even years? What if I told you there's a way to cut that timeline by up to 6x while guaranteeing accuracy? Datafold's Migration Agent is the only AI-powered solution that doesn't just translate your code; it validates every single data point to ensure perfect parity between your old and new systems. Whether you're moving from Oracle to Snowflake, migrating stored procedures to dbt, or handling complex multi-system migrations, they deliver production-ready code with a guaranteed timeline and fixed price. Stop burning budget on endless consulting hours. Visit dataengineeringpodcast.com/datafold to book a demo and see how they're turning months-long migration nightmares into week-long success stories.
    • Your host is Tobias Macey and today I'm interviewing Andy Warfield about S3 Tables and Vectors
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what your goals are with the Tables and Vector features of S3?
    • How did the experience of building S3 Tables inform your work on S3 Vectors?
    • There are numerous implementations of vector storage and search. How do you view the role of S3 in the context of that ecosystem?
    • The most directly analogous implementation that I'm aware of is the Lance table format. How would you compare the implementation and capabilities of Lance with what you are building with S3 Vectors?
      • What opportunity do you see for being able to offer a protocol compatible implementation similar to the Iceberg compatibility that you provide with S3 Tables?
    • Can you describe the technical implementation of the Vectors functionality in S3?
      • What are the sources of inspiration that you looked to in designing the service?
    • Can you describe some of the ways that S3 Vectors might be integrated into a typical AI application?
    • What are the most interesting, innovative, or unexpected ways that you have seen S3 Tables/Vectors used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on S3 Tables/Vectors?
    • When is S3 the wrong choice for Iceberg or Vector implementations?
    • What do you have planned for the future of S3 Tables and Vectors?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • S3 Tables
    • S3 Vectors
    • S3 Express
    • Parquet
    • Iceberg
    • Vector Index
    • Vector Database
    • pgvector
    • Embedding Model
    • Retrieval Augmented Generation
    • TwelveLabs
    • Amazon Bedrock
    • Iceberg REST Catalog
    • Log-Structured Merge Tree
    • S3 Metadata
    • Sentence Transformer
    • Spark
    • Trino
    • Daft
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    51 min
  • Revolutionizing Python Notebooks with Marimo
    Summary
    In this episode of the Data Engineering Podcast Akshay Agrawal from Marimo discusses the innovative new Python notebook environment, which offers a reactive execution model, full Python integration, and built-in UI elements to enhance the interactive computing experience. He discusses the challenges of traditional Jupyter notebooks, such as hidden states and lack of interactivity, and how Marimo addresses these issues with features like reactive execution and Python-native file formats. Akshay also explores the broader landscape of programmatic notebooks, comparing Marimo to other tools like Jupyter, Streamlit, and Hex, highlighting its unique approach to creating data apps directly from notebooks and eliminating the need for separate app development. The conversation delves into the technical architecture of Marimo, its community-driven development, and future plans, including a commercial offering and enhanced AI integration, emphasizing Marimo's role in bridging the gap between data exploration and production-ready applications.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Tired of data migrations that drag on for months or even years? What if I told you there's a way to cut that timeline by up to 6x while guaranteeing accuracy? Datafold's Migration Agent is the only AI-powered solution that doesn't just translate your code; it validates every single data point to ensure perfect parity between your old and new systems. Whether you're moving from Oracle to Snowflake, migrating stored procedures to dbt, or handling complex multi-system migrations, they deliver production-ready code with a guaranteed timeline and fixed price. Stop burning budget on endless consulting hours. Visit dataengineeringpodcast.com/datafold to book a demo and see how they're turning months-long migration nightmares into week-long success stories.
    • Your host is Tobias Macey and today I'm interviewing Akshay Agrawal about Marimo, a reusable and reproducible Python notebook environment
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you describe what Marimo is and the story behind it?
    • What are the core problems and use cases that you are focused on addressing with Marimo?
      • What are you explicitly not trying to solve for with Marimo?
    • Programmatic notebooks have been around for decades now. Jupyter was largely responsible for making them popular outside of academia. How have the applications of notebooks changed in recent years?
      • What are the limitations that have been most challenging to address in production contexts?
    • Jupyter has long had support for multi-language notebooks/notebook kernels. What is your opinion on the utility of that feature as a core concern of the notebook system?
    • Beyond notebooks, Streamlit and Hex have become quite popular for publishing the results of notebook-style analysis. How would you characterize the feature set of Marimo for those use cases?
    • For a typical data team that is working across data pipelines, business analytics, ML/AI engineering, etc. How do you see Marimo applied within and across those contexts?
    • One of the common difficulties with notebooks is that they are largely a single-player experience. They may connect into a shared compute cluster for scaling up execution (e.g. Ray, Dask, etc.). How does Marimo address the situation where a data platform team wants to offer notebooks as a service to reduce the friction to getting started with analyzing data in a warehouse/lakehouse context?
    • How are you seeing teams integrate Marimo with orchestrators (e.g. Dagster, Airflow, Prefect)?
    • What are some of the most interesting or complex engineering challenges that you have had to address while building and evolving Marimo?\
    • What are the most interesting, innovative, or unexpected ways that you have seen Marimo used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Marimo?
    • When is Marimo the wrong choice?
    • What do you have planned for the future of Marimo?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Marimo
    • Jupyter
    • IPython
    • Streamlit
      • Podcast.__init__ Episode
    • Vector Embeddings
    • Dimensionality Reduction
    • Kaggle
    • Pytest
    • PEP 723 script dependency metadata
    • MatLab
    • Visicalc
    • Mathematica
    • RMarkdown
    • RShiny
    • Elixir Livebook
    • Databricks Notebooks
    • Papermill
    • Pluto - Julia Notebook
    • Hex
    • Directed Acyclic Graph (DAG)
    • Sumble Kaggle founder Anthony Goldblum's startup
    • Ray
    • Dask
    • Jupytext
    • nbdev
    • DuckDB
      • Podcast Episode
    • Iceberg
    • Superset
    • jupyter-marimo-proxy
    • JupyterHub
    • Binder
    • Nix
    • AnyWidget
    • Jupyter Widgets
    • Matplotlib
    • Altair
    • Plotly
    • DataFusion
    • Polars
    • MotherDuck
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    52 min
  • Warehouse Native Incremental Data Processing With Dynamic Tables And Delayed View Semantics
    Summary
    In this episode of the Data Engineering Podcast Dan Sotolongo from Snowflake talks about the complexities of incremental data processing in warehouse environments. Dan discusses the challenges of handling continuously evolving datasets and the importance of incremental data processing for optimized resource use and reduced latency. He explains how delayed view semantics can address these challenges by maintaining up-to-date results with minimal work, leveraging Snowflake's dynamic tables feature. The conversation also explores the broader landscape of data processing, comparing batch and streaming systems, and highlights the trade-offs between them. Dan emphasizes the need for a unified theoretical framework to discuss semantic guarantees in data pipelines and introduces the concept of delayed view semantics, touching on the limitations of current systems and the potential of dynamic tables to simplify complex data workflows.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Dan Sotolongo about the challenges of incremental data processing in warehouse environments and how delayed view semantics help to address the problem
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by defining the scope of the term "incremental data processing"?
      • What are some of the common solutions that data engineers build when creating workflows to implement that pattern?
      • What are some common difficulties that they encounter in the pursuit of incremental data?
    • Can you describe what delayed view semantics are and the story behind it?
      • What are the problems that DVS explicitly doesn't address?
    • How does the approach that you have taken in Dynamic View Semantics compare to systems like Materialize, Feldera, etc.
    • Can you describe the technical architecture of the implementation of Dynamic Tables?
      • What are the elements of the problem that are as-yet unsolved?
      • How has the implementation changed/evolved as you learned more about the solution space?
    • What would be involved in implementing the delayed view semantics pattern in other dbms engines?
    • For someone who wants to use DVS/Dyamic Tables for managing their incremental data loads, what does the workflow look like?
      • What are the options for being able to apply tests/validation logic to a dynamic table while it is operating?
    • What are the most interesting, innovative, or unexpected ways that you have seen Dynamic Tables used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on Dynamic Tables/Delayed View Semantics?
    • When are Dynamic Tables/DVS the wrong choice?
    • What do you have planned for the future of Dynamic Tables?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Delayed View Semantics: Presentation Slides
    • Snowflake
    • NumPy
    • IPython
    • Jupyter
    • Flink
    • Spark Streaming
    • Kafka
    • Snowflake Dynamic Tables
    • Airflow
    • Dagster
    • Streaming Watermarks
    • Materialize
    • Feldera
    • ACID
    • CAP Theorem)
    • Linearizability
    • Serializable Consistency
    • SIGMOD
    • Materialized Views
    • dbt
    • Data Vault
    • Apache Iceberg
    • Databricks Delta
    • Hudi
    • Dead Letter Queue
    • pg_ivm
    • Property Based Testing
    • Iceberg V3 Row Lineage
    • Prometheus
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    56 min
  • Streamlining Data Pipelines with MCP Servers and Vector Engines
    Summary
    In this episode of the Data Engineering Podcast Kacper Łukawski from Qdrant about integrating MCP servers with vector databases to process unstructured data. Kacper shares his experience in data engineering, from building big data pipelines in the automotive industry to leveraging large language models (LLMs) for transforming unstructured datasets into valuable assets. He discusses the challenges of building data pipelines for unstructured data and how vector databases facilitate semantic search and retrieval-augmented generation (RAG) applications. Kacper delves into the intricacies of vector storage and search, including metadata and contextual elements, and explores the evolution of vector engines beyond RAG to applications like semantic search and anomaly detection. The conversation covers the role of Model Context Protocol (MCP) servers in simplifying data integration and retrieval processes, highlighting the need for experimentation and evaluation when adopting LLMs, and offering practical advice on optimizing vector search costs and fine-tuning embedding models for improved search quality.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • Your host is Tobias Macey and today I'm interviewing Kacper Łukawski about how MCP servers can be paired with vector databases to streamline processing of unstructured data
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • LLMs are enabling the derivation of useful data assets from unstructured sources. What are the challenges that teams face in building the pipelines to support that work?
    • How has the role of vector engines grown or evolved in the past ~2 years as LLMs have gained broader adoption?
      • Beyond its role as a store of context for agents, RAG, etc. what other applications are common for vector databaes?
    • In the ecosystem of vector engines, what are the distinctive elements of Qdrant?
    • How has the MCP specification simplified the work of processing unstructured data?
    • Can you describe the toolchain and workflow involved in building a data pipeline that leverages an MCP for generating embeddings?
    • helping data engineers gain confidence in non-deterministic workflows
    • bringing application/ML/data teams into collaboration for determining the impact of e.g. chunking strategies, embedding model selection, etc.
    • What are the most interesting, innovative, or unexpected ways that you have seen MCP and Qdrant used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on vector use cases?
    • When is MCP and/or Qdrant the wrong choice?
    • What do you have planned for the future of MCP with Qdrant?
    Contact Info
    • LinkedIn
    • Twitter/X
    • Personal website
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Qdrant
    • Kafka
    • Apache Oozi
    • Named Entity Recognition
    • GraphRAG
    • pgvector
    • Elasticsearch
    • Apache Lucene
    • OpenSearch
    • BM25
    • Semantic Search
    • MCP == Model Context Protocol
    • Anthropic Contextualized Chunking
    • Cohere
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    53 min
  • Foundational Data Engineering At Two Sigma
    Summary
    In this episode of the Data Engineering Podcast Effie Baram, a leader in foundational data engineering at Two Sigma, talks about the complexities and innovations in data engineering within the finance sector. She discusses the critical role of data at Two Sigma, balancing data quality with delivery speed, and the socio-technical challenges of building a foundational data platform that supports research and operational needs while maintaining regulatory compliance and data quality. Effie also shares insights into treating data as code, leveraging modern data warehouses, and the evolving role of data engineers in a rapidly changing technological landscape.


    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • This episode is brought to you by Coresignal, your go-to source for high-quality public web data to power best-in-class AI products. Instead of spending time collecting, cleaning, and enriching data in-house, use ready-made multi-source B2B data that can be smoothly integrated into your systems via APIs or as datasets. With over 3 billion data records from 15+ online sources, Coresignal delivers high-quality data on companies, employees, and jobs. It is powering decision-making for more than 700 companies across AI, investment, HR tech, sales tech, and market intelligence industries. A founding member of the Ethical Web Data Collection Initiative, Coresignal stands out not only for its data quality but also for its commitment to responsible data collection practices. Recognized as the top data provider by Datarade for two consecutive years, Coresignal is the go-to partner for those who need fresh, accurate, and ethically sourced B2B data at scale. Discover how Coresignal's data can enhance your AI platforms. Visit dataengineeringpodcast.com/coresignal to start your free 14-day trial. 
    • Your host is Tobias Macey and today I'm interviewing Effie Baram about data engineering in the finance sector
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by outlining the role of data in the context of Two Sigma?
    • What are some of the key characteristics of the types of data sources that you work with?
    • Your role is leading "foundational data engineering" at Two Sigma. Can you unpack that title and how it shapes the ways that you think about what you build?
      • How does the concept of "foundational data" influence the ways that the business thinks about the organizational patterns around data?
    • Given the regulatory environment around finance, how does that impact the ways that you think about the "what" and "how" of the data that you deliver to data consumers?
    • Being the foundational team for data use at Two Sigma, how have you approached the design and architecture of your technical systems?
      • How do you think about the boundaries between your responsibilities and the rest of the organization?
    • What are the design patterns that you have found most helpful in empowering data consumers to build on top of your work?
    • What are some of the elements of sociotechnical friction that have been most challenging to address?
    • What are the most interesting, innovative, or unexpected ways that you have seen the ideas around "foundational data" applied in your organization?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working with financial data?
    • When is a foundational data team the wrong approach?
    • What do you have planned for the future of your platform design?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • 2Sigma
    • Reliability Engineering
    • SLA == Service-Level Agreement
    • Airflow
    • Parquet File Format
    • BigQuery
    • Snowflake
    • dbt
    • Gemini Assist
    • MCP == Model Context Protocol
    • dtrace
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    56 min
  • Enabling Agents In The Enterprise With A Platform Approach
    Summary
    In this episode of the Data Engineering Podcast Arun Joseph talks about developing and implementing agent platforms to empower businesses with agentic capabilities. From leading AI engineering at Deutsche Telekom to his current entrepreneurial venture focused on multi-agent systems, Arun shares insights on building agentic systems at an organizational scale, highlighting the importance of robust models, data connectivity, and orchestration loops. Listen in as he discusses the challenges of managing data context and cost in large-scale agent systems, the need for a unified context management platform to prevent data silos, and the potential for open-source projects like LMOS to provide a foundational substrate for agentic use cases that can transform enterprise architectures by enabling more efficient data management and decision-making processes.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • This episode is brought to you by Coresignal, your go-to source for high-quality public web data to power best-in-class AI products. Instead of spending time collecting, cleaning, and enriching data in-house, use ready-made multi-source B2B data that can be smoothly integrated into your systems via APIs or as datasets. With over 3 billion data records from 15+ online sources, Coresignal delivers high-quality data on companies, employees, and jobs. It is powering decision-making for more than 700 companies across AI, investment, HR tech, sales tech, and market intelligence industries. A founding member of the Ethical Web Data Collection Initiative, Coresignal stands out not only for its data quality but also for its commitment to responsible data collection practices. Recognized as the top data provider by Datarade for two consecutive years, Coresignal is the go-to partner for those who need fresh, accurate, and ethically sourced B2B data at scale. Discover how Coresignal's data can enhance your AI platforms. Visit dataengineeringpodcast.com/coresignal to start your free 14-day trial. 
    • Your host is Tobias Macey and today I'm interviewing Arun Joseph about building an agent platform to empower the business to adopt agentic capabilities
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by giving an overview of how Deutsche Telekom has been approaching applications of generative AI?
      • What are the key challenges that have slowed adoption/implementation?
    • Enabling non-engineering teams to define and manage AI agents in production is a challenging goal. From a data engineering perspective, what does the abstraction layer for these teams look like? 
      • How do you manage the underlying data pipelines, versioning of agents, and monitoring of these user-defined agents?
    • What was your process for developing the architecture and interfaces for what ultimately became the LMOS?
      • How do the principles of operatings systems help with managing the abstractions and composability of the framework?
    • Can you describe the overall architecture of the LMOS?
      • What does a typical workflow look like for someone who wants to build a new agent use case?
      • How do you handle data discovery and embedding generation to avoid unnecessary duplication of processing?
    • With your focus on openness and local control, how do you see your work complementing projects like Oumi
    • What are the most interesting, innovative, or unexpected ways that you have seen LMOS used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on LMOS?
    • When is LMOS the wrong choice?
    • What do you have planned for the future of LMOS and MASAIC?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • LMOS
    • Deutsche Telekom
    • MASAIC
    • OpenAI Agents SDK
    • RAG == Retrieval Augmented Generation
    • LangChain
    • Marvin Minsky
    • Vector Database
    • MCP == Model Context Protocol
    • A2A (Agent to Agent) Protocol
    • Qdrant
    • LlamaIndex
    • DVC == Data Version Control
    • Kubernetes
    • Kotlin
    • Istio
    • Xerox PARC)
    • OODA (Observe, Orient, Decide, Act) Loop
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    55 min
  • Dagster's New Era: Modularizing Data Transformation in the Age of AI
    Summary
    In this episode of the Data Engineering Podcast we welcome back Nick Schrock, CTO and founder of Dagster Labs, to discuss the evolving landscape of data engineering in the age of AI. As AI begins to impact data platforms and the role of data engineers, Nick shares his insights on how it will ultimately enhance productivity and expand software engineering's scope. He delves into the current state of AI adoption, the importance of maintaining core data engineering principles, and the need for human oversight when leveraging AI tools effectively. Nick also introduces Dagster's new components feature, designed to modularize and standardize data transformation processes, making it easier for teams to collaborate and integrate AI into their workflows. Join in to explore the future of data engineering, the potential for AI to abstract away complexity, and the importance of open standards in preventing walled gardens in the tech industry.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • This episode is brought to you by Coresignal, your go-to source for high-quality public web data to power best-in-class AI products. Instead of spending time collecting, cleaning, and enriching data in-house, use ready-made multi-source B2B data that can be smoothly integrated into your systems via APIs or as datasets. With over 3 billion data records from 15+ online sources, Coresignal delivers high-quality data on companies, employees, and jobs. It is powering decision-making for more than 700 companies across AI, investment, HR tech, sales tech, and market intelligence industries. A founding member of the Ethical Web Data Collection Initiative, Coresignal stands out not only for its data quality but also for its commitment to responsible data collection practices. Recognized as the top data provider by Datarade for two consecutive years, Coresignal is the go-to partner for those who need fresh, accurate, and ethically sourced B2B data at scale. Discover how Coresignal's data can enhance your AI platforms. Visit dataengineeringpodcast.com/coresignal to start your free 14-day trial. 
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details. 
    • This is a pharmaceutical Ad for Soda Data Quality. Do you suffer from chronic dashboard distrust? Are broken pipelines and silent schema changes wreaking havoc on your analytics? You may be experiencing symptoms of Undiagnosed Data Quality Syndrome — also known as UDQS. Ask your data team about Soda. With Soda Metrics Observability, you can track the health of your KPIs and metrics across the business — automatically detecting anomalies before your CEO does. It’s 70% more accurate than industry benchmarks, and the fastest in the category, analyzing 1.1 billion rows in just 64 seconds. And with Collaborative Data Contracts, engineers and business can finally agree on what “done” looks like — so you can stop fighting over column names, and start trusting your data again.Whether you’re a data engineer, analytics lead, or just someone who cries when a dashboard flatlines, Soda may be right for you. Side effects of implementing Soda may include: Increased trust in your metrics, reduced late-night Slack emergencies, spontaneous high-fives across departments, fewer meetings and less back-and-forth with business stakeholders, and in rare cases, a newfound love of data. Sign up today to get a chance to win a $1000+ custom mechanical keyboard. Visit dataengineeringpodcast.com/soda to sign up and follow Soda’s launch week. It starts June 9th.
    • Your host is Tobias Macey and today I'm interviewing Nick Schrock about lowering the barrier to entry for data platform consumers
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by giving your summary of the impact that the tidal wave of AI has had on data platforms and data teams?
    • For anyone who hasn't heard of Dagster, can you give a quick summary of the project?
      • What are the notable changes in the Dagster project in the past year?
      • What are the ecosystem pressures that have shaped the ways that you think about the features and trajectory of Dagster as a project/product/community?
    • In your recent release you introduced "components", which is a substantial change in how you enable teams to collaborate on data problems. What was the motivating factor in that work and how does it change the ways that organizations engage with their data?
    • tension between being flexible and extensible vs. opinionated and constrained
    • increased dependency on orchestration with LLM use cases
    • reducing the barrier to contribution for data platform/pipelines
      • bringing application engineers into the mix
    • challenges of meeting users/teams where they are (languages, platform investments, etc.)
    • What are the most interesting, innovative, or unexpected ways that you have seen teams applying the Components pattern?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on the latest iterations of Dagster?
    • When is Dagster the wrong choice?
    • What do you have planned for the future of Dagster?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Links
    • Dagster+ Episode
    • Dagster Components Slide Deck
    • The Rise Of Medium Code
    • Lakehouse Architecture
    • Iceberg
    • Dagster Components
    • Pydantic Models
    • Kubernetes
    • Dagster Pipes
    • Ruby on Rails
    • dbt
    • Sling
    • Fivetran
    • Temporal
    • MCP == Model Context Protocol
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    1 hr 2 min
  • AI and the Lakehouse: How Starburst is Pioneering New Workflows
    Summary
    In this episode of the Data Engineering Podcast Alex Albu, tech lead for AI initiatives at Starburst, talks about integrating AI workloads with the lakehouse architecture. From his software engineering roots to leading data engineering efforts, Alex shares insights on enhancing Starburst's platform to support AI applications, including an AI agent for data exploration and using AI for metadata enrichment and workload optimization. He discusses the challenges of integrating AI with data systems, innovations like SQL functions for AI tasks and vector databases, and the limitations of traditional architectures in handling AI workloads. Alex also shares his vision for the future of Starburst, including support for new data formats and AI-driven data exploration tools.

    Announcements
    • Hello and welcome to the Data Engineering Podcast, the show about modern data management
    • Data migrations are brutal. They drag on for months—sometimes years—burning through resources and crushing team morale. Datafold's AI-powered Migration Agent changes all that. Their unique combination of AI code translation and automated data validation has helped companies complete migrations up to 10 times faster than manual approaches. And they're so confident in their solution, they'll actually guarantee your timeline in writing. Ready to turn your year-long migration into weeks? Visit dataengineeringpodcast.com/datafold today for the details.
    • This is a pharmaceutical Ad for Soda Data Quality. Do you suffer from chronic dashboard distrust? Are broken pipelines and silent schema changes wreaking havoc on your analytics? You may be experiencing symptoms of Undiagnosed Data Quality Syndrome — also known as UDQS. Ask your data team about Soda. With Soda Metrics Observability, you can track the health of your KPIs and metrics across the business — automatically detecting anomalies before your CEO does. It’s 70% more accurate than industry benchmarks, and the fastest in the category, analyzing 1.1 billion rows in just 64 seconds. And with Collaborative Data Contracts, engineers and business can finally agree on what “done” looks like — so you can stop fighting over column names, and start trusting your data again.Whether you’re a data engineer, analytics lead, or just someone who cries when a dashboard flatlines, Soda may be right for you. Side effects of implementing Soda may include: Increased trust in your metrics, reduced late-night Slack emergencies, spontaneous high-fives across departments, fewer meetings and less back-and-forth with business stakeholders, and in rare cases, a newfound love of data. Sign up today to get a chance to win a $1000+ custom mechanical keyboard. Visit dataengineeringpodcast.com/soda to sign up and follow Soda’s launch week. It starts June 9th. This episode is brought to you by Coresignal, your go-to source for high-quality public web data to power best-in-class AI products. Instead of spending time collecting, cleaning, and enriching data in-house, use ready-made multi-source B2B data that can be smoothly integrated into your systems via APIs or as datasets. With over 3 billion data records from 15+ online sources, Coresignal delivers high-quality data on companies, employees, and jobs. It is powering decision-making for more than 700 companies across AI, investment, HR tech, sales tech, and market intelligence industries. A founding member of the Ethical Web Data Collection Initiative, Coresignal stands out not only for its data quality but also for its commitment to responsible data collection practices. Recognized as the top data provider by Datarade for two consecutive years, Coresignal is the go-to partner for those who need fresh, accurate, and ethically sourced B2B data at scale. Discover how Coresignal's data can enhance your AI platforms. Visit dataengineeringpodcast.com/coresignal to start your free 14-day trial.
    • Your host is Tobias Macey and today I'm interviewing Alex Albu about how Starburst is extending the lakehouse to support AI workloads
    Interview
    • Introduction
    • How did you get involved in the area of data management?
    • Can you start by outlining the interaction points of AI with the types of data workflows that you are supporting with Starburst?
    • What are some of the limitations of warehouse and lakehouse systems when it comes to supporting AI systems?
    • What are the points of friction for engineers who are trying to employ LLMs in the work of maintaining a lakehouse environment?
    • Methods such as tool use (exemplified by MCP) are a means of bolting on AI models to systems like Trino. What are some of the ways that is insufficient or cumbersome?
    • Can you describe the technical implementation of the AI-oriented features that you have incorporated into the Starburst platform?
      • What are the foundational architectural modifications that you had to make to enable those capabilities?
    • For the vector storage and indexing, what modifications did you have to make to iceberg?
      • What was your reasoning for not using a format like Lance?
    • For teams who are using Starburst and your new AI features, what are some examples of the workflows that they can expect?
    • What new capabilities are enabled by virtue of embedding AI features into the interface to the lakehouse?
    • What are the most interesting, innovative, or unexpected ways that you have seen Starburst AI features used?
    • What are the most interesting, unexpected, or challenging lessons that you have learned while working on AI features for Starburst?
    • When is Starburst/lakehouse the wrong choice for a given AI use case?
    • What do you have planned for the future of AI on Starburst?
    Contact Info
    • LinkedIn
    Parting Question
    • From your perspective, what is the biggest gap in the tooling or technology for data management today?
    Closing Announcements
    • Thank you for listening! Don't forget to check out our other shows. Podcast.__init__ covers the Python language, its community, and the innovative ways it is being used. The AI Engineering Podcast is your guide to the fast-moving world of building AI systems.
    • Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes.
    • If you've learned something or tried out a project from the show then tell us about it! Email [email protected] with your story.
    Links
    • Starburst
      • Podcast Episode
    • AWS Athena
    • MCP == Model Context Protocol
    • LLM Tool Use
    • Vector Embeddings
    • RAG == Retrieval Augmented Generation
      • AI Engineering Podcast Episode
    • Starburst Data Products
    • Lance
    • LanceDB
    • Parquet
    • ORC
    • pgvector
    • Starburst Icehouse
    The intro and outro music is from The Hug by The Freak Fandango Orchestra / CC BY-SA
    45 min

About Data Engineering Podcast

From the publisher's feed

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some…

More shows like Data Engineering Podcast

This Week in Startups by Jason Calacanis

This Week in Startups

1,290 Listeners

The Changelog: Software Development, Open Source by Changelog Media

The Changelog: Software Development, Open Source

286 Listeners

The a16z Show by Andreessen Horowitz

The a16z Show

1,087 Listeners

Software Engineering Daily by Software Engineering Daily

Software Engineering Daily

623 Listeners

Risky Business by Risky Business Media

Risky Business

375 Listeners

Talk Python To Me by Michael Kennedy

Talk Python To Me

582 Listeners

Super Data Science: ML & AI Podcast with Jon Krohn by Jon Krohn

Super Data Science: ML & AI Podcast with Jon Krohn

305 Listeners

NVIDIA AI Podcast by NVIDIA

NVIDIA AI Podcast

338 Listeners

Syntax - Tasty Web Development Treats by Wes Bos & Scott Tolinski - Full Stack JavaScript Web Developers

Syntax - Tasty Web Development Treats

985 Listeners

Practical AI by Daniel Whitenack and Chris Benson

Practical AI

203 Listeners

Dwarkesh Podcast by Dwarkesh Patel

Dwarkesh Podcast

565 Listeners

The Data Engineering Show by The Firebolt Data Bros

The Data Engineering Show

8 Listeners

Latent Space: The AI Engineer Podcast by Latent.Space

Latent Space: The AI Engineer Podcast

102 Listeners

This Day in AI Podcast by Michael Sharkey, Chris Sharkey

This Day in AI Podcast

222 Listeners

The AI Daily Brief: Artificial Intelligence News and Analysis by Nathaniel Whittemore

The AI Daily Brief: Artificial Intelligence News and Analysis

684 Listeners