Data Mesh Radio

Data Mesh Radio

By Data as a Product Podcast NetworkNewsTechnologyEducationTech News
Download on the App Store

Data Mesh Radio episodes

  • #53 To Analytics Engineer or Not to Analytics Engineer - That's a Question - Mesh Musings 9

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    A mesh musings asking if we should embed data/analytics engineers into the domains to serve as the data product developers or not.

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/

    All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf

    14 min
  • #52 Data Mesh Data Governance: Getting Out of Your Own Way - Interview w/ Sarita Bakst

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.

    Data Mesh Learning Meetup Presentation: https://www.youtube.com/watch?v=7iazNKG8XQo

    Sarita's LinkedIn: https://www.linkedin.com/in/saritabakst/

    In this episode, Scott interviewed Sarita Bakst, Managing Director and a leader in Firm-wide Data Management at JP Morgan Chase. Sarita was previously on a Data Mesh Learning meetup and is helping lead the firm's data mesh governance charge.

    Per Sarita, historically, data governance has meant controls and gatekeeping to most people - typically getting in way of innovation. So there needs to be a focus on changing that narrative, not just through words but actions to show it's not the case. In data mesh, you need to ensure that domains can make good decisions on governance and seek out subject matter experts when it makes sense.

    Sarita covered that one of the key issues in the way governance has been done is the people making decisions - the central governance team - don't have the real understanding of the data. When those decisions are put in the hands of the people who really know the data, but with guardrails and guidance, the fear is lifted about can we actually use this data and how. This opens up lots of new opportunities to leverage your data.

    Sarita strongly recommends starting with purpose-built data products. Find a use case and build data products to serve that specific. And data products MUST be about unlocking business value. You don't need to serve up all of a domain's data on day one, in version one of that domain's first data product - make it extendible and reusable so you can find additional consumers and expand over time.

    Get out of your own way on data governance in data mesh. You are going to learn and your approach will evolve as you learn. It's okay to not know everything upfront, set yourself up to not get in trouble - put the proper guardrails in place - but you won't know everything. Think of designing your risk controls as toll-gates and make sure they aren't bottlenecks.

    Have standards (not standardization) so people don't have to invent things from scratch. Standards for interoperability, naming, etc. - think of them as guiding...

    1 hr 11 min
  • #51 A DevOps Angle to Data Mesh and WePay's Journey - Interview w/ Chris Riccomini

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.

    What the Heck is a Data Mesh?! Post: https://cnr.sh/essays/what-the-heck-data-mesh

    Chris' Twitter: https://twitter.com/criccomini

    Chris' LinkedIn: https://www.linkedin.com/in/riccomini/

    Chris' website: https://cnr.sh/

    The Missing README book: https://themissingreadme.com/

    In this episode, Scott interviewed Chris Riccomini, a Software Engineer, Author, and Investor. Chris led the infrastructure team at WePay when they embarked on a data mesh journey and made a well-written post on thinking about data mesh in DevOps terms.

    Like a number of people/organizations that have come on the podcast, at WePay, Chris was pursuing the general goals of data mesh and was applying some of the approaches as well - but it was not nearly as cohesive as Zhamak laid things out.

    Their initial setup had two teams managing the pipeline/transformation infrastructure. Chris's team was mostly handing the extracting and loading and then there was a team of analytics engineers handling the transformations. The Transformation team saw a major increase in demand and quickly became overloaded -> a bottleneck. Chris' team also started to get overloaded so they knew they had to evolve.

    One way the team started to address the bottlenecks was by decentralizing the pipelines. Teams could make a request and a scalable and reliable pipeline would essentially get automatically set up for them. WePay is in the financial services space so as part of those pipelines, to prevent risk, teams could mark their sensitive/PII columns and the infra team also put in some autodetection capabilities to make sure they didn't miss any.

    WePay created a "canonical data representation" or CDR, which is pretty analogous to a data product in data mesh. Chris really liked WePay's use of the embedded analytics engineer to serve as a data product developer.

    One key innovation for WePay was tooling to...

    1 hr 16 min
  • Weekly Episode Summaries and Programming Notes - Week of Apr 4, 2022 - Data Mesh Radio

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/

    All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf

    29 min
  • #50 Evaluating if Data Mesh is a Fit; Team Structures; and WTF is Federated Computational Governance - Interview w/ Marius Ingjer

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here

    Marius' LinkedIn: https://www.linkedin.com/in/marius-ingjer-a155313/

    Marius' Medium: https://medium.com/@marius.ingjer

    In this episode, Scott interviewed Marius Ingjer, Co-founder and Senior Consultant at Knirkefritt AS. Marius has been working with a few clients on implementing data mesh.

    They covered 3 distinct topics: evaluating if data mesh is a fit for your organization, team structure challenges in data mesh or data mesh-like implementations, and a simplified definition of federated computational governance.

    Evaluating if data mesh is right for you:

    To start, Marius provided a list of evaluation questions to help you determine if data mesh might be right for you:

    1. How many data sets are you producing?
    2. What is the lead time to creating a new dataset?
    3. How well are your datasets serving your data needs?
    4. How many domains do you have?
    5. How complex are your domains?
    6. How does the team respond to new data requirements?
    7. How usable in general is your data?

    Every company wants to share data well but the centralized data team isn't the bottleneck yet for many. Centralization can add a lot of value until it starts to become more hurtful than helpful and yes, figuring out that point is easier said than done. Centralization of data fights Conway's Law and can become way too much cognitive load so it will eventually become an issue for many organizations.

    A key question in evaluating if data mesh is a fit: what is the cost of allowing your data processes to fail? Per Marius, the business consequence of failed reports has historically not been that high. But if you are driving business decisions, whether that is ML or just crucial day-to-day decisions on your data, data mesh might become more attractive.

    Team structures and challenges in data mesh:

    In general, it's important to understand that implementing data mesh will cause cultural challenges - Marius believes developers generally don't want to ALSO share their data. It's additional work so you have to align incentives, which is far easier said than done.

    That...

    1 hr 19 min
  • #49 Differentiating the Baby from the Bathwater in Data Mesh - Mesh Musings 8

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    An episode about how to figure out what is good to reuse from your existing data approach and what to toss out when it comes to data mesh.

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/

    All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf

    13 min
  • #48 Overcoming Obstinate Organizational Obstacles in Data Mesh - Interview w/ Scott Hawkins

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here

    Scott Hawkins' LinkedIn: https://www.linkedin.com/in/scott-hawkins-8934393/

    In this episode, Scott interviewed Scott Hawkins, Principal Data Architect at ITV.

    Scott views data mesh as a mechanism for change. Your company culture and your understanding of it are crucial to establishing data mesh well, driving that buy-in. Ask yourself: what challenges does data mesh actually address and hopefully solve, how will it impact the business not the tech, and what does it change. Your organization might not be ready for data mesh. Or a specific domain might not be ready. And that's okay!

    As ITV moved forward, they found a "good enough" solution via a global ID. It's not perfect, there might be some overlap - such as one person might have a different global ID for their online subscription versus their broadcast subscription - but it is far better than what they were doing. And it allows for interoperability/joins across the data. This is a big improvement - don't let perfect be the enemy of good or done.

    One thing working for ITV is deploying a "team-in-a-box" to help domains move forward - similar to an internal consulting team. Each situation is different so each box they are given is different. The team-in-a-box concept also means it is somewhat easier to build common best practices internally. Coming to the table with defaults has really helped ITV.

    Per Scott, there are 3 good ways to drive buy-in for the domain teams:

    1. At the senior level - so it trickles down as the management for the team is bought in.
    2. Via a strong carrot - solve a problem for them as a kind of quid quo pro / mutually beneficial solution. Trying to solve an unrelated problem will drive lower buy-in.
    3. Work on realigning the team KPIs/OKRs with the senior leaders to actually realign incentives.

    Continuing on the driving buy-in, Scott recommends working with the domain managers to generate a viable/valuable carrot for the entire team. Explain to those leaders why it matters, work with the leaders to revamp the KPIs if the KPIs are getting in the way of delivering a good data product. This is why exec-level buy-in is so crucial - it is pretty hard to start modifying team KPIs/OKRs without it!...

    57 min
  • #47 Skipping the Fluff of Domain Driven Design for Data Mesh - Interview w/ Lorenzo Nicora

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here

    Lorenzo's LinkedIn: https://www.linkedin.com/in/nicus/

    Pat Helland Data on the Inside/Outside: http://cidrdb.org/cidr2005/papers/P12.pdf

    Mesh-AI careers page: https://www.mesh-ai.com/join-us

    In this episode, Scott interviewed Lorenzo Nicora, Principal Data Consultant at data mesh and AI focused consultancy Mesh-AI.

    Scott asked Lorenzo to be on to continue the series of interviews on domain driven design (or DDD) for data. It is a topic that many are struggling with so having lots of perspectives on it is crucial. On the episode title, a key output was explicit permission to skip a lot of the tactical patterns of DDD. Others have also said similar things but I wanted to make sure it was explicit.

    Before we jump into the DDD parts, Lorenzo made a good point on your data mesh Proof of Concept / starting your journey. You need to start with manageable problems. Start with a consumer-driven problem but a source/producer-aligned data product. There is a lot of nuance in the interview on why this matters.

    Per Lorenzo, identifying the domains is crucial but it is the hardest part of DDD. That shouldn't scare you because you can start with things being a bit blurry. It's important to understand your high-level domains but you can get moving without mapping out all of your domains.

    A key theme from Lorenzo: the language is at the center of everything in DDD. It is part of the data modeling and it goes all the way down to the code.

    Per Lorenzo, DDD is all about communication, knowledge capture, and knowledge sharing. Knowledge capture is about extracting knowledge and then writing it down. Knowledge sharing is about finding scalable ways to share context.

    Some advice/pointers from Lorenzo:

    1. Teams have to truly understand the language of their own domain - remove the ambiguities, even if that feels like it's putting in too much work.
    2. Event storming is a great way to approach tackling DDD for Data.
    3. Event sourcing is crucial for modeling the problem of the domain.
    4. Terminology is very key...
    1 hr 22 min
  • Weekly Episode Summaries and Programming Notes - Week of Mar 27, 2022 - Data Mesh Radio

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    If you want to learn more and/or join the Data Mesh Learning Community, see here: https://datameshlearning.com/community/

    All music used this episode was found on PixaBay and was created by (including slight edits by Scott Hirleman): Lesfm, MondayHopes, SergeQuadrado, ItsWatR, Lexin_Music, and/or nevesf

    28 min
  • #46 Designing a Data Literacy Approach for Data Engineers - Interview w/ Dan Sullivan

    Sign up for Data Mesh Understanding's free roundtable and introduction programs here: https://landing.datameshunderstanding.com/

    Please Rate and Review us on your podcast app of choice!

    If you want to be a guest or give feedback (suggestions for topics, comments, etc.), please see here

    Episode list and links to all available episode transcripts here.

    Provided as a free resource by Data Mesh Understanding / Scott Hirleman. Get in touch with Scott on LinkedIn if you want to chat data mesh.

    Transcript for this episode (link) provided by Starburst. See their Data Mesh Summit recordings here and their great data mesh resource center here.

    Dan's LinkedIn: https://www.linkedin.com/in/dansullivanpdx/

    Dan's Email: dan.sullivan at 4mile.io

    In this episode, Scott interviewed Dan Sullivan, Principal Data Architect at 4 Mile Analytics.

    A key point Dan brought up is tech debt around data. Taking on tech debt should ALWAYS be a very conscious choice. But the way most organizations work with data, it is much more of an unconscious choice, especially by data producers, who are taking on debt that the data engineering teams will have to pay down. We need to find ways to deliver value quickly but with discipline.

    Zhamak has mentioned in a few talks that data engineers soon may not exist in orgs deploying data mesh. Dan actually somewhat agrees that data engineering will change a lot as right now, there is a big rush to build out the initial iterations of data products (the industry definition). Going forward, Dan thinks there will be a need for data engineers that can really understand consumer needs and build the interactions, e.g. the SDKs, to leverage data.

    Dan has 3 key pillars for driving data literacy for data engineers are domain knowledge, learning, and collaboration. Data engineers should pair with business people to acquire domain knowledge, they should be given the opportunity to spend time doing things like online training to learn, and they should collaborate across the organization instead of just being ticket tacklers.

    Per Dan, not all data engineers are the same depending on background - some come from a data analyst/data science background but many come from a software engineering background. So we can't treat training all data engineers as if it's the same. But we do need them to have a well-rounded background. A big need is for them to understand more about the data consumers and/or the producers so embedding them in the domains can really help.

    For driving buy-in with data engineers, Dan points to the problems typically being around incentives. Data engineering is often hampered by organizational issues and a lack of clear direction. So if you can tackle those, you can often win over DEs.

    In any organization but especially in one implementing data mesh, standards, protocols, and contracts are all very important....

    1 hr 1 min

About Data Mesh Radio

From the publisher's feed

Interviews with data mesh practitioners, deep dives/how-tos, anti-patterns, panels, chats (not debates) with skeptics, "mesh musings", and so much more. Host Scott Hirleman (founder of the Data Mesh Learning Community) shares his learnings - and those of the broader data community - from over a year of deep diving into data mesh.