
Sign up to save your podcasts
Or


Yaser Najafi helps decide which data startups Snowflake partners with and puts money behind: 70+ investments since 2019. His filter isn't whether a company uses AI. It's whether it understands the workflow that existed before the model. He explains most enterprise AI efforts fail on data rather than models, why agents don't care which platform your data sits on, and why context without governance is only half the problem. Plus: what CoCo changes about who gets to touch data, what comes after it, and the one problem that will get a founder a meeting with Snowflake Ventures.
β± CHAPTERS 0:00 Welcome to Data Splash 1:25 The 30-Second Splash 2:20 From alliances to writing checks 3:32 Why being a builder changes how you evaluate startups 4:38 What Snowflake Ventures is actually funding right now 6:23 Vertical solutions, and why data is the defensible moat 7:58 Unstructured data and the context layer 8:25 Agents don't care what platform the data lives on 10:55 The governance layer cake: can you, and should you 12:26 CoCo: what Cortex Code changes 15:47 Who gets to touch data now 18:37 Using CoCo before it went public 20:50 What comes after CoCo: cloud agents and shared skills 22:42 Why enterprises trust an in-platform agent 23:19 The biggest opportunity for a startup today 25:32 Documents, video, audio: the pipeline problem 26:13 Wrap and the Silicon Valley AI Hub
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Yaser Najafi: [https://www.linkedin.com/in/ynajafi/] β’ Connect with Ido Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#Snowflake #AgenticAI #DataEngineering #EnterpriseAI #AIAgents
Every company wants a company brain. Jessica Talisman β semantic engineer, creator of the Ontology Pipeline, and a librarian before she built knowledge graphs at Adobe and Amazon β explains why most of them are doing it backwards. She calls it epistemic squatting: assuming you can lift and shift data practices into a discipline shaped over 3,000 years. Inside: why your database holds a fraction of what your organization knows, why "tacit knowledge" is the most misused phrase in the field, and where to actually start.
β± CHAPTERS 0:00 Introduction 1:23 The 30 Second Splash 2:09 From Spielberg's Shoah Foundation to enterprise knowledge graphs 3:42 Why "move fast and break things" fails at knowledge work 5:59 What knowledge actually is and why context isn't the same thing 8:19 Taxonomies, SKOS, and defining things so machines can use them 9:32 Advice for data engineers: think beyond your database 13:50 The "company brain" and epistemic squatting 17:22 Three years of ungoverned AI documents, and the EU AI Act 18:45 Symbolic AI, GOFAI, and the neurosymbolic revival 21:19 Augmentation before automation, and capturing tacit knowledge 23:51 The Ontology Pipeline book: where organizations should start 25:32 Closing
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Jessica Talisman: [https://www.linkedin.com/in/jmtalisman/] β’ Connect with Ido Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#KnowledgeEngineering #AI #DataEngineering #LLMs #AIAgents
Tiankai Feng became a governance leader because he wouldn't stop complaining about bad data. Completely understandable! He breaks down the three root causes of data quality problems - technical, process, and human - and why the human ones are nearly impossible to fix downstream. Plus what AI changed: Models need far more historical data than governance ever scoped, and GenAI shifted focus to unstructured documents nobody has governed. Also: the shared drive with fifty versions of the same PDF, and why use cases beat cleanup projects. Hosted by Ido Bronstein of Upriver.
β± CHAPTERS 0:00 Introduction 1:06 The 30-Second Splash 1:56 Meet Tiankai: career path and the two books 3:35 Why he left analytics for governance 6:51 Why data quality is so hard: intended vs. actual data use 8:42 Three root causes: technical, process, and human error 11:06 What maturity looks like: from reactive to proactive 13:24 Why use cases beat cleanup projects 16:23 How AI changed the governance mandate 19:44 Governing unstructured data 21:36 How AI is making stewardship leaner 24:05 The data governance songs 24:49 Advice for new stewards 26:11 Closing takeaways
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Tiankai Feng: [https://www.linkedin.com/in/tiankaifeng/] β’ Connect with Ido Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#DataGovernance #AI #DataEngineering #LLMs #AIAgents
Teams keep attaching AI agents to their CRM, then realize they need to deduplicate it first. Abe Gong, CEO and co-founder of Great Expectations, joins Ido Bronstein to explain why organizational memory β not the model β is the real bottleneck for AI. They cover why "context graph" is the most overhyped phrase in data, the curation test every knowledge base should pass, how agents ended the analyst bottleneck, and why data teams must evolve from technical function to the organization's arbiter of truth. Plus: why a good knowledge base makes company culture legible for the first time.
β± CHAPTERS 0:00 Welcome and guest intro 1:02 The 30 Second Splash 1:45 Abe's path from algorithms to engineering problems 4:03 What actually counts as a knowledge base 6:12 Curation: why storage alone isn't enough 8:24 How companies manage knowledge today 10:04 The landscape: search, semantic layers, and skills 12:31 Shifting bottlenecks: code review and the end of the analyst gate 14:23 Nine definitions of churn β democratization, solved? 15:30 How knowledge bases reshape the data team 18:49 The data team as the organization's judge 19:52 Planning for agents that do the technical work 22:04 The prize: humanβagent collaboration 23:34 Your knowledge base is your culture 26:13 Takeaways and wrap
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Abe Gong: [https://www.linkedin.com/in/abe-gong-8a77034/] β’ Connect with Ido Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#DataEngineering #AI #DataPlatform #LLMs #AIAgents
For a decade, the data moat was infrastructureβwhoever could afford 20 engineers to maintain scrapers won. Nimble's Amaury Desrosiers explains why that moat is gone, how external web data became a first-class citizen alongside your warehouse, and why real-time will soon be a default property, not a category. Plus the three-circle model for internal vs. external data and the one move every data leader should make tomorrow.
β± CHAPTERS 0:00 Introduction 0:51 The 30 Second Splash 1:44 What Nimble does 2:08 Questions internal data can't answer 3:22 The three concentric circles 4:28 Ground truth vs. context β joining the two 5:40 Why web data used to be a nightmare 6:27 Solving connection and parsing end to end 7:53 The moat moves from infra to usage 9:32 Why real-time value compounds 11:55 Batch vs. ad hoc pipelines 14:24 One layer, not one connector per site 15:34 Schema by use case 16:58 Scale changed, the data model didn't 18:58 Three predictions for the next three years 20:25 What data leaders should do tomorrow 21:19 Takeaways and wrap
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Amaury Desrosiers: [https://www.linkedin.com/in/amaurydesrosiers] β’ Connect with Omri Lifshitz: [https://www.linkedin.com/in/omri-lifshitz-8a531814a/]
#DataEngineering #AI #DataPlatform #LLMs #AIAgents
Twenty years of new tools were supposed to make data engineering easier. Shachar Meir (former Director of Data Engineering at Meta, former data lead at PayPal) argues they made it more chaotic. He explains how "store now, model later" broke the data contract we're now scrambling to rebuild, why data teams fail for reasons that have nothing to do with technology, and the two tips he gives every CDO. Plus the photographer analogy that explains why he's still not convinced data engineering is going anywhere. Hosted by Ido Bronstein on Data Splash.
β± CHAPTERS 0:00 Introduction 0:59 The 30 Second Splash 1:34 What a data advisor actually does 2:50 Why data teams fail: the missing ingredients 4:02 The DBA era β when data was always modeled 7:03 Breaking the contract: Hadoop and data lakes 8:45 Did technology make the job easier? 9:46 Two tips for every CDO 11:34 Enter AI: risks and the how-vs-what shift 15:03 Shifting ownership to the business 16:46 Is managing data still a profession? 18:46 The photographer analogy 19:06 Career advice for early-career data people 21:34 Why the business matters more than ever 22:28 Wrap-up
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Shachar Meir: [https://www.linkedin.com/in/shacharmeir/] β’ Connect with Ido Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#DataEngineering #AI #DataPlatform #LLMs #AIAgents
Everyone says they need data lineage. Few can say why. Harel Shein, Senior Engineering Manager at Datadog and an OpenLineage steering committee member, makes the case that AI turns lineage from a nice-to-have graph into infrastructure. We cover what OpenLineage actually is, why agents querying your warehouse need it to understand what data means, the maintainer tax of vibe-coded PRs, and Harel's wishlist: push the standard into every engine like Spark, and extend it to ML. One spec to rule them all.
β± CHAPTERS 0:00 Intro β What we're getting into 1:05 Rapid-fire questions 2:18 What OpenLineage actually is 3:22 Why Apache Airflow built it in (and the Linux Foundation) 5:10 How AI is reshaping open source maintenance 6:10 Vibe-coded PRs, maintainer burden, and the AGENTS.md fix 7:37 Why coding agents understand open source better than closed source 8:35 Real-world use cases: ops, data quality, compliance, cost 11:38 How AI changes the lineage use case β amplification 13:59 Lineage as the foundation for agents 16:05 MCP, self-serve data, and the trust problem 19:37 Lineage for unstructured data β the hard problem 22:03 If you had an army of engineers, what would you build? 24:39 Where data engineering is heading 26:29 Takeaways and one spec to rule them all
π ABOUT DATA SPLASH: Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: https://www.upriverdata.com/ β’ Connect with Harel Shein: https://www.linkedin.com/in/harelshein/ β’ Connect with Ido Bronstein: https://www.linkedin.com/in/ido-bronstein/
#DataEngineering #AI #Datadog #LLMs #AIAgents
Atlassian's Head of Data Engineering & AI Enablement, Prakash Reddy, joins host Ido Bronstein, Upriver's Co-founder and CEO, for an honest conversation about using AI to automate data engineering work at scale.
Six months in, Atlassian is shipping real results, 5-day tickets in under 3 days, an on-call agent that triages production failures, and a clear roadmap for AI-ready data. But the most useful part of this episode is what had to be true before any of it worked: a multi-year migration to declarative YAML pipelines, environment isolation, and a clean medallion architecture.
In this episode: β’ Why AI doesn't work without foundational data architecture β’ The 3 pillars Atlassian picked for AI ROI (and what they skipped) β’ How to measure productivity gains in total cost of ownership, not velocity β’ Why hallucinations (3-4 out of 10) are a workflow problem, not a model problem β’ The "coalition of the willing" approach, bottom-up experiments + top-down consolidation β’ A prediction on role convergence: knowledge engineer, context engineer, agent orchestrator β’ Why the moat stops being SQL, and what replaces it
Whether you're a data engineer trying to make sense of where the role is headed, a data leader planning an AI rollout, or just curious how a company at Atlassian's scale is approaching this, this episode is built for you.
β± CHAPTERS 00:00 Intro & 30-Second Splash 02:00 How Atlassian's data org is structured 05:00 Why automate data engineering with AI? 08:00 The foundation that made AI possible 10:00 The 3 pillars: incremental dev, on-call, AI-ready data 13:00 Real productivity numbers (and how to measure them honestly) 17:00 Hallucinations, guardrails, and what actually breaks 21:00 Org design: bottom-up + top-down 25:00 The future of data roles β convergence is coming 30:00 Closing thoughts
π ABOUT DATA SPLASH Data Splash is a podcast for data engineers, data leaders, and anyone trying to make sense of AI and data right now. Brought to you by Upriver.
π Subscribe for new episodes weekly.
π LINKS β’ Upriver: [https://www.upriverdata.com/] β’ Connect with Prakash Reddy: [https://www.linkedin.com/in/prakashreddy1357/] β’ Connect with IdoΒ
Bronstein: [https://www.linkedin.com/in/ido-bronstein/]
#DataEngineering #AI #Atlassian #DataPlatform #LLMs #AIAgents
From the publisher's feed
An Al Data Engineering Podcast