
Sign up to save your podcasts
Or


Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Justin's LinkedIn: https://www.linkedin.com/in/justincinmd/
In this episode, Scott interviewed Justin Cunningham, who worked as a tech lead and data architect on data platforms at Netflix, Yelp, and Atlassian over the last 8.5 years. In that time, Justin was involved in initiatives to push data ownership to developers / domains.
To sum up one of Justin's points he touched on repeatedly - he recommends to create a pool of low effort data which will inherently have low quality. Use that for initial research into what might be useful. Focus on maximizing accessibility - you can have governance and use things like join restrictions or give consumers an ability to self-certify that they are using the data responsibly. Once you get the use cases, then you go for the data mesh quality data products.
Justin saw a lot of success at Yelp focusing on data availability - getting data to a place it could be found and played with - was a bigger driver for success than focusing initially on data quality. Once people discovered what data was available and how they might use it, the organization was able to work towards getting that data to an acceptable quality level.
Another point Justin made was figure out which you want to optimize for in general: getting things right upfront or testing and changing. He believes in optimizing for change. Create an adaptive process and optimize for learning. Keep it simple and focus on value delivery - it will set up more tractable bets.
At Yelp, they were trying to ETL a huge amount of data in their data warehouse to build reports for the C-Suite. But they were never really going to get enough data ingested to really meet their goals. It was taking them 2 weeks to create each new set of ETLs and that was just creation, not maintenance - it was looking like they'd need 5x the number of people.
What Justin found the most useful at Yelp was to focus on getting as much "usable" data in an automated way. They achieved this initially through the data mesh anti-pattern of copying direct from the underlying operational data stores and building business logic on top of it. But, that data getting into the hands of the data team meant there could be an initial value assessment - once they proved...
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Doron's LinkedIn: https://www.linkedin.com/in/porat-doron/
Our journey towards an open data platform: https://medium.com/yotpoengineering/our-journey-towards-an-open-data-platform-8cfac98ef9f5
A simplified, lightweight ETL Framework based on Apache Spark by Yotpo: https://github.com/YotpoLtd/metorikku
The Data Swamp (in Hebrew): https://open.spotify.com/show/5YDdtRhp1RVw7r5fbYFtPQ?si=5x4HzOyhTX6n46HqY5kV6w&nd=1
In this episode, Scott interviewed Doron Porat, a Data Infrastructure Leader at the SaaS company Yotpo.
Some crucial points Doron made:
1) Be kind to yourself when you make mistakes - it's worse to stagnate so don't be afraid of change and making choices
2) Build versus buy is always tough but don't let your ego get in the way and push you towards building everything
3) If you do buy, build a close relationship with your vendors to help influence the roadmap and have an outlet if you are having issues
4) A data platform team's job is to drive usage as usage means creating value - drive towards that and set your KPIs around platform usage
5) There will likely be many different types of consumers of your data platform - work to improve / optimize the user experience for most folks
Doron is a technologist at heart so for each decision she instinctually wants to build instead of buy. And at the start of building out the data platform for Yotpo, that was typically her decision. But as the demands for more and more capabilities from the platform, the increasing ubiquity and quality/scalability of as-a-service offerings, and the growing need to drive usage and developer happiness instead of manage cool tech, she started to consume more and more managed services.
When you are building out the platform, vendors, in the long run, can often better serve your needs because they have a whole lot...
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.
Samia's LinkedIn: https://www.linkedin.com/in/samia-rahman-b7b65216/
FHIR standard cheat sheet: https://www.healthit.gov/topic/standards-technology/standards/fhir-fact-sheets
In this episode, Scott interviewed Samia Rahman, Director of Data and AI Strategy and Architecture at life sciences company Seagen. Samia is helping to lead Seagen's early data mesh implementation after helping with two implementations at Thoughtworks since the start of 2019.
For Samia, interoperability is about taking information from two systems and combining them to get a higher value. A simple definition but a good one.
Two potential key takeaways:
1) don't try to plan too much ahead for developing interoperability standards but definitely keep an eye out for places where you could start to develop those standards. And your standards really, really should evolve - you don't have to nail them right out of the gate.
2) your interoperability will also evolve - you don't need to make every data product interoperable with every other data product and you can start with basic interoperability first. The more you can standardize around unique identifiers, the better, but it's okay to not get it right first thing out of the gate.
Samia started her career - and even before in school - focusing on software, especially end-to-end development. A repeating pattern for her has been how crucial contract testing is to getting things into a trustable and scalable state. We've had them in hardware and software for a long time and if you don't have easy testing, those systems often get replaced pretty quickly. Those tests are the safety net to allow for fast and reliable evolution. And that evolution is a key theme for this conversation - set yourself up to iterate and evolve as you learn. Work to not paint yourself in a corner
Data standards, including specifically for interoperability, are everywhere in the life sciences space - FHIR, FDA has lots, etc. but it's still not great for truly sharing the meaning of the data. FAIR is trying to get there but the interoperability and domain...
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Scott shares his views on the importance of collaboration via negotiation, not requests, to make your data mesh implementation a success.
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Abe's Twitter: @AbeGong / https://twitter.com/AbeGong
Abe's LinkedIn: https://www.linkedin.com/in/abe-gong-8a77034/
Great Expectations Community Page: https://greatexpectations.io/community
In this episode, Scott interviewed Abe Gong, the co-creator Great Expectations (an open source data quality / monitoring / observability tool) and co-founder/CEO of Superconductive.
One caveat before jumping in is that Abe is passionate about the topic and has created tooling to help address it. So try to view Abe's discussion of Great Expectations as an approach rather than a commercial for the project/product.
To start the conversation, Abe shared some of his background experience living the pain of unexpected upstream data changes causing data chaos / lots of work to recover from and adapt. Part of where we need to get to using something like data contracts is to remove the need to recover in addition to adapting and move towards controlled/expected adaptation. Abe believes that the best framing for data contracts is to think about them as a set of expectations.
To define expectations here, this would include not just schema but also the content of data, such as value ranges/types/distributions/relationships across tables/etc. So for instance, a column may be a one to five for rankings and then the application team changes it one to 10. The schema may not be broken - it is still passing whole numbers - but the new range is not within expectations so the contract is broken.
At current, Abe sees the best way to not break social expectations is via getting consumers and producers in a meeting to talk about the upcoming changes and prepare, such as with versioning. But, as tooling improves, Abe sees a world where we won't even need a lot of those meetings going forward - either because data pipelines can be "self-healing" and automatically adapt to changes upstream or because metadata and tools for context-sharing will reduce the need for meetings.
Abe sees two distinct use cases in general for data contracts or more specifically
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here
Sadie's LinkedIn: https://www.linkedin.com/in/sadie-martin-06404125/
In this episode, Scott interviewed Sadie Martin, Senior Product Manager, Data Platform at Q4 Inc about applying a product mindset to data in general. This is really crucial to getting data as a product right but also in building out your data platforms and even some processes for data mesh.
Scott's summation of some key points:
Sadie started as a data analyst where the team didn't have a product manager - they were doing a lot of work and weren't sure if things were likely to work or even if what they did had a positive impact after it was done. So she started to take on some of the task of answering those questions and transitioned into being a product manager for data.
So, what is a product mindset? For Sadie, the easy definition but with lots of hidden depth, is "it's all about really understanding the problem". For most organizations, really thinking about the problem you are trying to solve is new relative to data. There may be a data request but what product or process is that data contributing to and what is that product or process trying to solve?
Sadie believes measuring the problem is really crucial. Once you figure out what you are trying to solve, what is the scope of the problem? How are you going to measure if you are actually solving the problem? Especially is it better than what you were previously doing? She also talked about the importance of customer-centricity - really why are they making a data ask? Should this really be a one-off or a repeatable process? Did they ask for the complete set of what they need? Etc.
One crucial insight Sadie has brought from product management to data is to be willing and ready to throw things...
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.
This episode is part of the Data Innovation Summit Takeover week of Data Mesh Radio.
Data Innovation Summit website: https://datainnovationsummit.com/; use code DATAMESHR20G for 20% off tickets
Free Ticket Raffle for Data Innovation Summit (submissions must be by April 25 at 11:59pm PST): Google Form
Henrik's LinkedIn: https://www.linkedin.com/in/henrikgothberg/
Dairdux website: https://dairdux.com/
Airplane Alliance website: https://airplanealliance.com/
In the last of the interviews for the Data Innovation Summit Takeover week, Scott interviewed Henrik Göthberg, the Founder and CEO of consulting company Dairdux, the Co-Founder of the Airplane Alliance, and the Chairman of the Data Innovation Summit.
Let's start with some conclusions/advice from Henrik:
Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/
Please Rate and Review us on your podcast app of choice!
If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here
Episode list and links to all available episode transcripts here.
Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.
If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/
All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf
From the publisher's feed