
Sign up to save your podcasts
Or


In this episode of Data Engineering Weekly Radio, we delve into modern data stacks under pressure and the potential consolidation of the data industry. We refer to a four-part article series that explores the data infrastructure landscape and the Software as a Service (SaaS) products available in data engineering, machine learning, and artificial intelligence.
We discussed that the siloed nature of many data products has led to industry consolidation, ultimately benefiting customers. Throughout our discussion, we touch on how the Modern Data Stack (MDS) movement has resulted in various specialized tools in areas such as ingestion, cataloging, governance, and quality. However, we also acknowledge that as budgets tighten and CFOs become more cautious, the market is now experiencing a push toward bundling and consolidation.
In this consolidation, we explore the roles of large players like Snowflake, Databricks, and Microsoft and cloud companies like AWS and Google. We debate who will be the "control center" of the data workload, as many companies claim to be the central component in the data ecosystem. As hosts, we agree it's difficult to predict the industry's future, but we anticipate the market will mature and settle soon.
We discussed the potential consolidation of various tools and categories in the modern data stack, including ETL, reverse ETL, data quality, observability, and data catalogs. Consolidation is likely, as many of these tools share common ground and can benefit from unified experiences for users. We also explored how tools like DBT, Airflow, and Databricks could emit information about data lineage, potentially leading to a "catalog of catalogs" that centralizes the visualization and governance of data.
We suggested that the convergence of data quality, observability, and catalogs would revolve around ensuring clean, trusted data that is easily discoverable. We also touched on the role of data lineage and pondered whether the control of data lineage would translate to control over the entire data stack. We considered the possibility that orchestration engines might step into data quality, observability, and catalogs, leading to further consolidation in the industry.
We also acknowledged the shift in conversation within the data community from focusing on technology comparisons to examining organizational landscapes and the production and consumption of data. We agreed that there is still much room for innovation in this space and that consolidating features is more beneficial than competing with one another.
We contemplated how tools like DBT might extend their capabilities by tackling other aspects of the data stack, such as ingestion. Additionally, we discussed the potential consolidation in the MLOps space, with various tools stepping on each other's territory as they address customer needs.
Overall, we emphasized the importance of unifying user experiences and blurring the lines between individual categories in the data infrastructure landscape. We also noted the parallels between feature stores and data products, suggesting that there may be further convergence between MLOps and data engineering practices in the future. Ultimately, customer delight and experience are the driving forces behind these developments.
We also discussed ETL's potential future, the rise of zero ETL, and its challenges. Additionally, we touched on the growing importance of data products and contracts, emphasizing the need for a contract-first approach in building successful data products.
In conclusion, Matt Turck's blog provided us with an excellent opportunity to discuss and analyze the current trends in the data industry. We look forward to seeing how these trends continue to evolve and shape the future of data management and analytics. Until the next edition, take care, and see you all!
Reference
https://mattturck.com/mad2023/
https://mattturck.com/mad2023-part-iii/
In this episode of Data Engineering Weekly Radio, we delve into modern data stacks under pressure and the potential consolidation of the data industry. We refer to a four-part article series that explores the data infrastructure landscape and the Software as a Service (SaaS) products available in data engineering, machine learning, and artificial intelligence.
We discussed that the siloed nature of many data products has led to industry consolidation, ultimately benefiting customers. Throughout our discussion, we touch on how the Modern Data Stack (MDS) movement has resulted in various specialized tools in areas such as ingestion, cataloging, governance, and quality. However, we also acknowledge that as budgets tighten and CFOs become more cautious, the market is now experiencing a push toward bundling and consolidation.
In this consolidation, we explore the roles of large players like Snowflake, Databricks, and Microsoft and cloud companies like AWS and Google. We debate who will be the "control center" of the data workload, as many companies claim to be the central component in the data ecosystem. As hosts, we agree it's difficult to predict the industry's future, but we anticipate the market will mature and settle soon.
We discussed the potential consolidation of various tools and categories in the modern data stack, including ETL, reverse ETL, data quality, observability, and data catalogs. Consolidation is likely, as many of these tools share common ground and can benefit from unified experiences for users. We also explored how tools like DBT, Airflow, and Databricks could emit information about data lineage, potentially leading to a "catalog of catalogs" that centralizes the visualization and governance of data.
We suggested that the convergence of data quality, observability, and catalogs would revolve around ensuring clean, trusted data that is easily discoverable. We also touched on the role of data lineage and pondered whether the control of data lineage would translate to control over the entire data stack. We considered the possibility that orchestration engines might step into data quality, observability, and catalogs, leading to further consolidation in the industry.
We also acknowledged the shift in conversation within the data community from focusing on technology comparisons to examining organizational landscapes and the production and consumption of data. We agreed that there is still much room for innovation in this space and that consolidating features is more beneficial than competing with one another.
We contemplated how tools like DBT might extend their capabilities by tackling other aspects of the data stack, such as ingestion. Additionally, we discussed the potential consolidation in the MLOps space, with various tools stepping on each other's territory as they address customer needs.
Overall, we emphasized the importance of unifying user experiences and blurring the lines between individual categories in the data infrastructure landscape. We also noted the parallels between feature stores and data products, suggesting that there may be further convergence between MLOps and data engineering practices in the future. Ultimately, customer delight and experience are the driving forces behind these developments.
We also discussed ETL's potential future, the rise of zero ETL, and its challenges. Additionally, we touched on the growing importance of data products and contracts, emphasizing the need for a contract-first approach in building successful data products.
We also shared our thoughts on the potential convergence of various categories, like data cataloging and data contracts, which could give rise to more comprehensive and powerful data solutions. Furthermore, we discussed the significance of interfaces and their potential to shape the future of the data stack.
In conclusion, Matt Turck's blog provided us with an excellent opportunity to discuss and analyze the current trends in the data industry. We look forward to seeing how these trends continue to evolve and shape the future of data management and analytics. Until the next edition, take care, and see you all!
Reference
https://mattturck.com/mad2023/
https://mattturck.com/mad2023-part-iii/
Subscribe to www.dataengineeringweekly.com
From Data Engineering Weekly Edition #121, we took the following articles
Oda: Data as a product at Oda
Oda writes an exciting blog about “Data as a Product,” describing why we must treat data as a product, dashboard as a product, and the ownership model for data products.
https://medium.com/oda-product-tech/data-as-a-product-at-oda-fda97695e820
The blog highlights six key principles of the value creation of data.
https://medium.com/oda-product-tech/the-six-principles-for-how-we-run-data-insight-at-oda-ba7185b5af39
Ashwin & Ananth Conversation Highlights
Peter Bruins: Some reflections on talking with Data leaders
Data Mesh/ Data Product/ Data Contract all the concepts trying to address this problem, and this is a Billion $ $ $ worth of a problem to solve. The author leaves a bigger question, Ownership plays a central role in all these concepts, but what is the incentive to bring Ownership?
https://www.linkedin.com/pulse/some-reflections-talking-data-leaders-peter-bruins/
Ashwin & Ananth Conversation Highlights
Faire: The great migration from Redshift to Snowflake
Is Redshift dying? I’m seeing an increasing pattern of people migrating from Redshift to Snowflake or Lakehouse. Flair wrote a detailed blog on the reasoning behind Redshift to Snowflake migration, its journey, and its key takeaway.
https://craft.faire.com/the-great-migration-from-redshift-to-snowflake-173c1fb59a52
Flair also opensource some of the utility scripts to make your life easier to move from Redshift to Snowflake
https://github.com/Faire/snowflake-migration
Ashwin & Ananth Conversation Highlights
We thank all the writes of the blog for sharing their knowledge with the data community
We are back in our Data Engineering Weekly Radio for edition #121. We will take 2 or 3 articles from each week's Data Engineering Weekly edition and go through an in-depth analysis.
Please subscribe to our Podcast on your favorite apps.
From editor #121, we took the following articles
Oda: Data as a product at Oda
Oda writes an exciting blog about “Data as a Product,” describing why we must treat data as a product, dashboard as a product, and the ownership model for data products.
https://medium.com/oda-product-tech/data-as-a-product-at-oda-fda97695e820
The blog highlights six key principles of the value creation of data.
* Domain knowledge + discipline expertise
* Distributed Data Ownership and shared Data Ownership
* Data as a Product
* Enablement over Handover
* Impact through Exploration and Experimentation
* Proactive attitude towards Data Privacy & Ethics
https://medium.com/oda-product-tech/the-six-principles-for-how-we-run-data-insight-at-oda-ba7185b5af39
Here are a few highlights from the podcast
"Oda builds the whole data product principle & the implementation structure being built on top of the core values, instead of reflecting any industry jargons.”
"Don't make me think. The moment you make your users think, you lose your value proposition as a platform or a product.”
"The platform enables the domain; domain enables your consumer. It's a chain of value creation going on top and like simplifying everyone's life, accessing data, making informed decisions.”
"I think putting that, documenting it, even at the start of it, I think that's where the equations start proving themselves. And that's essentially what product thinking is all about.”
Peter Bruins: Some reflections on talking with Data leaders
Data Mesh/ Data Product/ Data Contract all the concepts trying to address this problem, and this is a Billion $ $ $ worth of a problem to solve. The author leaves a bigger question, Ownership plays a central role in all these concepts, but what is the incentive to bring Ownership?
https://www.linkedin.com/pulse/some-reflections-talking-data-leaders-peter-bruins/
Here are a few highlights from the podcast
"Ownership. It's all about the ownership." - Peter Burns.
"The weight of the success (growth of adoption) of the data leads to its failure.
Faire: The great migration from Redshift to Snowflake
Is Redshift dying? I’m seeing an increasing pattern of people migrating from Redshift to Snowflake or Lakehouse. Flair wrote a detailed blog on the reasoning behind Redshift to Snowflake migration, its journey, and its key takeaway.
https://craft.faire.com/the-great-migration-from-redshift-to-snowflake-173c1fb59a52
Flair also opensource some of the utility scripts to make your life easier to move from Redshift to Snowflake
https://github.com/Faire/snowflake-migration
Here are a few highlights from the podcast
"If you left like one percent of my data is still in Redshift and 99% of your data in Snowflake, you're degrading your velocity and the quality of your delivery.”
Please read Data Engineering Weekly Edition #120
Topic 1: Colin Campbell: The Case for Data Contracts - Preventative data quality rather than reactive data quality
In this episode, we focus on the importance of data contracts in preventing data quality issues. We discuss an article by Colin Campbell highlighting the need for a data catalog and the market scope for data contract solutions. We also touch on the idea that data creation will be a decentralized process and the role of tools like data contracts in enabling successful decentralized data modeling. We emphasize the importance of creating high-quality data and the need for technological and organizational solutions to achieve this goal.
Key highlights of the conversation
Link:
https://uncomfortablyidiosyncratic.substack.com/p/the-case-for-data-contracts
https://www.dataengineeringweekly.com/p/introducing-schemata-a-decentralized
Topic 2: Yerachmiel Feltzman: Action-Position data quality assessment framework
In this conversation, we discuss a framework for data quality assessment called the Action Position framework. The framework helps define what actions should be taken based on the severity of the data quality problem. We also discuss two patterns for data quality: Write-Audit-Publish (WAP) and Audit-Write-Publish (AWP). The WAP pattern involves writing data, auditing it, and publishing it, while the AWP pattern involves auditing data, writing it, and publishing it. We encourage readers to share their best practices for addressing data quality issues.
Are you using any Data Quality framework in your organization? Do you have any best practices on how you address data quality issues? What do you think of the action-position data quality framework? Please add your comments in the SubStack chat.
Link:
https://medium.com/everything-full-stack/action-position-data-quality-assessment-framework-d833f6b77b7
Dremio WAP pattern: https://www.dremio.com/resources/webinars/the-write-audit-publish-pattern-via-apache-iceberg/
Topic 3: Guy Fighel - Stop emphasizing the Data Catalog
We discuss the limitations of data catalogs and the author’s view on the semantic layer as an alternative. The author argues that data catalogs are passive and quickly become outdated and that a stronger contract with enforced data quality could be a better solution. We also highlight the cost factors of implementing a data catalog and suggest that a more decentralized approach may be necessary to keep up with the increasing number of data sources. Innovation in this space is needed to improve organizations' discoverability and consumption of data assets.
Link:
https://www.linkedin.com/pulse/stop-emphasizing-data-catalog-guy-fighel/
https://www.dataengineeringweekly.com/p/data-catalog-a-broken-promise
We are back in our Data Engineering Weekly Radio for edition #120. We will take 2 or 3 articles from each week's Data Engineering Weekly edition and go through an in-depth analysis.
From editor #120, we took the following articles
Topic 1: Colin Campbell: The Case for Data Contracts - Preventative data quality rather than reactive data quality
In this episode, we focus on the importance of data contracts in preventing data quality issues. We discuss an article by Colin Campbell highlighting the need for a data catalog and the market scope for data contract solutions. We also touch on the idea that data creation will be a decentralized process and the role of tools like data contracts in enabling successful decentralized data modeling. We emphasize the importance of creating high-quality data and the need for technological and organizational solutions to achieve this goal.
Key highlights of the conversation
"Preventative data quality rather than reactive data quality. It should start with contracts." - Colin Campbell. - Author of the article
"Contracts put a preventive structure in place" - Ashwin.
"The successful data-driven companies all do one thing very well. They create high-quality data." - Ananth.
Ananth’s post on Schemata
Topic 2: Yerachmiel Feltzman: Action-Position data quality assessment framework
In this conversation, we discuss a framework for data quality assessment called the Action Position framework. The framework helps define what actions should be taken based on the severity of the data quality problem. We also discuss two patterns for data quality: Write-Audit-Publish (WAP) and Audit-Write-Publish (AWP). The WAP pattern involves writing data, auditing it, and publishing it, while the AWP pattern involves auditing data, writing it, and publishing it. We encourage readers to share their best practices for addressing data quality issues.
Are you using any Data Quality framework in your organization? Do you have any best practices on how you address data quality issues? What do you think of the action-position data quality framework? Please add your comments in the SubStack chat.
https://medium.com/everything-full-stack/action-position-data-quality-assessment-framework-d833f6b77b7
Dremio WAP pattern: https://www.dremio.com/resources/webinars/the-write-audit-publish-pattern-via-apache-iceberg/
Topic 3: Guy Fighel - Stop emphasizing the Data Catalog
We discuss the limitations of data catalogs and the author’s view on the semantic layer as an alternative. The author argues that data catalogs are passive and quickly become outdated and that a stronger contract with enforced data quality could be a better solution. We also highlight the cost factors of implementing a data catalog and suggest that a more decentralized approach may be necessary to keep up with the increasing number of data sources. Innovation in this space is needed to improve organizations' discoverability and consumption of data assets.
Something to think about in this conversation
"If you don't catalog everything and we only catalog what is required for the purpose of business decision-making, does that solve the data catalog problem in an organization?"
https://www.linkedin.com/pulse/stop-emphasizing-data-catalog-guy-fighel/
We are super excited to be back to discussing Data Engineering Weekly Newsletter articles every week. We will take 2 or 3 articles from each week's Data Engineering Weekly edition and go through an in-depth analysis.
On Data Engineering Weekly edition #119, We are taking three articles.
#1 Netflix's article about Scaling Media Machine Learning at Netflix
https://netflixtechblog.com/scaling-media-machine-learning-at-netflix-f19b400243
#2 Alex Woodie's article about Open Table Formats Square Off in Lakehouse Data Smackdown
https://www.datanami.com/2023/02/15/open-table-formats-square-off-in-lakehouse-data-smackdown/
#3 Plum Living's article about Building a semantic layer in Preset (Superset) with dbt
https://medium.com/plum-living/building-a-semantic-layer-in-preset-superset-with-dbt-71ee3238fc20
We referenced David Jayatillake's article about Metricalypse in the show.
We are super excited to be back to discussing Data Engineering Weekly Newsletter articles every week. We will take 2 or 3 articles from each week's Data Engineering Weekly edition and go through an in-depth analysis.
On Data Engineering Weekly edition #119, We are taking three articles.
https://netflixtechblog.com/scaling-media-machine-learning-at-netflix-f19b400243
https://www.datanami.com/2023/02/15/open-table-formats-square-off-in-lakehouse-data-smackdown/
https://medium.com/plum-living/building-a-semantic-layer-in-preset-superset-with-dbt-71ee3238fc20
We referenced David Jayatillake's article about Metricalypse in the show.
https://davidsj.substack.com/p/metricalypse-now
From the publisher's feed

274 Listeners

286 Listeners

622 Listeners

581 Listeners

304 Listeners

145 Listeners

226 Listeners

265 Listeners

959 Listeners

179 Listeners

204 Listeners

140 Listeners

514 Listeners

8 Listeners

102 Listeners