Data Mesh: A Deep Dive
Data mesh is a new way of thinking about how to manage data in large organizations. It is a decentralized, domain-oriented approach that aims to address the limitations of traditional centralized data models, such as scalability challenges, data silos, and slow response times to changing data needs.
Zhamak Dehghani: The Creator of Data Mesh
Zhamak Dehghani is the creator of the data mesh paradigm. She is the director of emerging technologies at Thoughtworks, a global software consultancy1. Dehghani has over 20 years of experience in the technology industry, with a focus on distributed systems and big data architecture. She is a passionate advocate for decentralized technology solutions, and she believes that data mesh is a key enabler of data democratization. Dehghani also contributes to the Thoughtworks Technology Advisory Board and the Thoughtworks Technology Radar. She has contributed to multiple patents in distributed computing communications, as well as embedded device technologies.
In 2018, Dehghani founded the concept of data mesh, a paradigm shift in big data management toward data decentralization. She is also the author of the book, Data Mesh: Delivering Data-Driven Value at Scale. In 2022, Dehghani left Thoughtworks to found Nextdata Technologies to focus on decentralized data.
What is Data Mesh?
Data mesh is a novel approach to enterprise data management that emphasizes decentralization and domain ownership to address the challenges of sharing, accessing, and managing analytical data in complex and large-scale environments. It is a distributed architecture that organizes data around business domains, creating self-contained data products. One way to think about this is that data mesh, like microservices, structures data around business domains, creating self-contained data products. This approach promotes data flexibility and interoperability, similar to how microservices enhance software development by breaking down monolithic applications into smaller, independent services.
Data mesh leverages the principles of domain-driven design, a software development paradigm that aligns the structure and language of code with its corresponding business domain, to deliver a self-serve data platform. This allows users to abstract the technical complexity and focus on their individual data use cases. In a data mesh, data ownership and responsibilities are distributed among domain-specific teams or data product teams, granting them autonomy in managing their data within their respective domains. This decentralized approach aims to address the limitations associated with centralized data models, such as scalability challenges, data silos, and slow response times to changing data needs.
Data mesh also addresses advanced data security challenges through distributed, decentralized ownership. Organizations have multiple data sources from different lines of business that must be integrated for analytics. A data mesh architecture effectively unites the disparate data sources and links them together through centrally managed data sharing and governance guidelines. Business functions can maintain control over how shared data is accessed, who accesses it, and in what formats it’s accessed.
Data mesh is a relatively new concept that gained significant traction after the pandemic. As organizations are actively experimenting with different technological approaches to build data meshes for specific use cases, it is clear that enterprise-wide implementation is still in its early stages.
It is important to understand that data mesh is not just a technical architecture, but also a socio-technical paradigm. This means that it requires changes to the way people, processes, and tools are organized within an organization. Data mesh advocates for bringing the operational and analytical worlds closer together by decentralizing and realigning the ownership of analytical data from a central team to domains.
To further illustrate the key differences between traditional data management platforms and data mesh architectures, consider the following table:
Principles of Data Mesh
Data mesh is based on four key principles:
- Domain-oriented ownership: Data is owned and managed by the teams that use it. This ensures that data is treated as a product and that it is managed in a way that meets the needs of the business. For example, in a retail company, the team responsible for managing customer data would also be responsible for managing the data related to customer purchases, returns, and interactions with customer service. This allows the team to have a holistic view of the customer and make better decisions about how to serve them.
- Data as a product: Data is treated as a product, with a clear lifecycle and defined quality standards. This ensures that data is discoverable, understandable, and trustworthy. Just like any other product, data needs to be designed, built, tested, and maintained. It also needs to meet certain quality standards in order to be useful. By treating data as a product, organizations can ensure that it is managed in a way that maximizes its value.
- Self-serve data infrastructure: A self-serve data infrastructure provides domain teams with the tools and capabilities they need to manage their data independently. This includes tools for data ingestion, processing, storage, and access. This allows domain teams to be more agile and responsive to changing business needs. They can access the data they need, when they need it, without having to rely on a central data team.
- Federated computational governance: Federated computational governance ensures that data is managed in a consistent and compliant way across the organization. This includes establishing global rules, naming conventions, and best practices for documentation. This ensures that data is consistent, accurate, and secure, regardless of where it is stored or who is using it.
Benefits of Data Mesh
Data mesh offers several benefits over traditional centralized data models, including:
- Increased agility: Data mesh enables domain teams to move faster and be more responsive to changing business needs. By decentralizing data ownership and management, data mesh empowers domain teams to make decisions more quickly and efficiently. This agility is essential in today’s rapidly changing business environment.
- Improved scalability: Data mesh can scale to meet the needs of even the largest organizations. As organizations grow and their data volumes increase, data mesh can easily accommodate the additional data and users. This scalability is essential for organizations that want to be able to adapt to future growth.
- Reduced costs: Data mesh can help to reduce costs by eliminating the need for a centralized data team. By distributing data management responsibilities to domain teams, organizations can reduce the overhead associated with a centralized data team. This can lead to significant cost savings. However, it is important to note that data mesh still requires some level of central governance and coordination.
- Enhanced data quality: Data mesh can help to improve data quality by ensuring that data is managed by the teams that use it. Domain teams are more likely to have a deep understanding of the data they are managing and are therefore more likely to be able to identify and correct data quality issues.
- Increased data democratization: Data mesh makes data more accessible to everyone in the organization, which can lead to better decision-making. By making data more readily available, data mesh empowers everyone in the organization to make data-driven decisions. This can lead to better business outcomes.
- Reduced Bottlenecks and Increased Efficiency: Data mesh can lead to fewer bottlenecks, more reliability, and faster results. By empowering domain teams to manage their own data, data mesh reduces the reliance on a central data team, which can often be a bottleneck for data access and analysis. This leads to more efficient data operations and faster time-to-insight.
- Improved Accessibility: Data mesh reduces bottleneck issues and makes data more easily accessible to all users across your organization. By organizing data by its domain, data mesh eliminates the need for users to navigate complex data pipelines and access data directly from the source. This improved accessibility can lead to faster decision-making and increased data literacy across the organization.
- Facilitated Self-Service: Data mesh facilitates self-service applications from multiple data sources, broadening the access of data beyond more technical resources, such as data scientists, data engineers, and developers. By making data more discoverable and accessible via this domain-driven design, it reduces data silos and operational bottlenecks, enabling faster decision-making and freeing up technical users to prioritize tasks that better utilize their skillsets.
Use Cases of Data Mesh
Data mesh can be used in a variety of use cases, including:
- Customer 360: Data mesh can be used to create a single view of the customer, which can be used to improve customer service and marketing. By combining data from various sources, such as sales, marketing, and customer service, organizations can gain a complete understanding of their customers and their needs. This can lead to more personalized and effective customer interactions.
- Fraud detection: Data mesh can be used to detect fraud by analyzing data from multiple sources. By analyzing data from different domains, such as transactions, customer behavior, and account activity, organizations can identify patterns that may indicate fraudulent activity.
- Supply chain optimization: Data mesh can be used to optimize the supply chain by tracking inventory levels and predicting demand. By having real-time access to data from different parts of the supply chain, organizations can make better decisions about inventory management, production planning, and logistics.
- Personalized recommendations: Data mesh can be used to provide personalized recommendations to customers. By analyzing customer data, such as purchase history, browsing behavior, and preferences, organizations can provide more relevant and targeted recommendations.
- Real-time analytics: Data mesh can be used to provide real-time insights into business operations. By having access to real-time data from different domains, organizations can monitor key performance indicators, identify trends, and make quick decisions to optimize their operations.
- Third-Party and Public Datasets: Data mesh can be applied to use cases that require third-party and public datasets. You can treat external data as a separate domain and implement it in the mesh to ensure consistency with internal datasets. This allows organizations to leverage external data sources to gain a more comprehensive understanding of their business and the market.
Data Mesh vs. Data Lake
While data mesh can leverage a data lake as its central data store, it is not a storage solution itself but a data management architecture. Data lakes are centralized repositories designed to store vast amounts of data in a scalable and cost-effective manner. Data mesh, on the other hand, decentralizes data ownership, distributing responsibilities to individual domains or business units, fostering a more collaborative and scalable approach.
Here’s a table summarizing the key differences:
| Feature | Data Mesh | Data Lake |
|---|---|---|
| Architecture | Decentralized | Centralized |
| Data Ownership | Domain-oriented | Centralized |
| Data Management | Self-serve | IT-managed |
| Focus | Data as a product | Data storage |
| Scalability | Highly scalable | Scalable |
| Agility | High | Moderate |
| Data Governance | Federated | Centralized |
Data Mesh vs. Data Warehouse
Data mesh and data warehouses can coexist and serve different purposes within an organization’s data infrastructure. Data warehouses are centralized repositories that store structured data that has been cleaned and transformed. They are often used for reporting and analysis. Data mesh, on the other hand, is a more decentralized and domain-oriented approach that empowers individual teams to manage their data as products.
Here’s a table summarizing the key differences:
| Feature | Data Mesh | Data Warehouse |
|---|---|---|
| Architecture | Decentralized | Centralized |
| Data Ownership | Domain-oriented | Centralized |
| Data Management | Self-serve | IT-managed |
| Data Structure | Variety of formats | Structured |
| Focus | Data as a product | Historical data analysis |
| Scalability | Highly scalable | Scalable |
| Agility | High | Moderate |
| Data Governance | Federated | Centralized |
When to Consider Data Mesh
While data mesh offers numerous benefits, it’s not a one-size-fits-all solution. Organizations should consider several factors when deciding whether to implement data mesh, including:
- Quantity of data sources: The more data sources an organization has, the more complex it becomes to manage data centrally. Data mesh can help to simplify data management by distributing responsibilities to domain teams.
- Size of your data team: If your data team is small and struggling to keep up with the demands of the business, data mesh can help to alleviate the burden by empowering domain teams to manage their own data.
- Number of data domains: The more functional teams (marketing, sales, operations, etc.) that need access to data, the more beneficial data mesh becomes. Data mesh allows each domain to manage its own data, making it easier for teams to access the data they need.
- Data engineering bottlenecks: If the data engineering team is frequently a bottleneck to the implementation of new data products, data mesh can help to remove this bottleneck by empowering domain teams to manage their own data pipelines.
- Data governance: Data mesh requires a strong data governance framework to ensure that data is managed in a consistent and compliant way across the organization.
Impetus Behind Data Mesh
While data lakes and data warehouses offer certain advantages, data mesh emerged as a response to the evolving challenges of managing data in complex organizations. Traditional centralized data models are often not able to keep up with the growing volume, velocity, and variety of data. Data mesh offers a more scalable and agile approach to data management.
Some of the specific challenges that data mesh is trying to address include:
- Data silos: Data silos occur when data is stored in different systems and is not accessible to everyone in the organization. Data mesh breaks down data silos by making data more discoverable and accessible.
- Data bottlenecks: Data bottlenecks occur when a centralized data team is responsible for managing all of the data in the organization. Data mesh eliminates data bottlenecks by empowering domain teams to manage their own data.
- Data quality issues: Data quality issues can occur when data is not managed in a consistent way across the organization. Data mesh addresses data quality issues by establishing global rules and standards for data management.
Companies Using Data Mesh
A number of companies are currently using data mesh, including:
- Zalando: Zalando is a German online fashion retailer that is using data mesh to improve customer experience and personalize recommendations.
- Netflix: Netflix is a streaming service that is using data mesh to improve its recommendation engine and personalize content for users.
- Intuit: Intuit is a financial software company that is using data mesh to improve its fraud detection capabilities.
- VistaPrint: VistaPrint is an online printing company that is using data mesh to improve its marketing campaigns.
- PayPal: PayPal is an online payments company that is using data mesh to improve its risk management capabilities.
Conclusion
Data mesh is a new paradigm for data management that offers several benefits over traditional centralized data models. It is a decentralized, domain-oriented approach that is well-suited for organizations that need to be more agile and responsive to changing business needs. A number of companies are already using data mesh to improve their business operations, and it is likely that data mesh will become increasingly popular in the years to come.
Data mesh has the potential to revolutionize the way organizations manage and use data. By breaking down data silos, empowering domain teams, and promoting data as a product, data mesh can help organizations to become more data-driven and agile. However, implementing data mesh is not without its challenges. Organizations need to be prepared to make changes to their organizational structure, culture, and technology.
As data mesh continues to evolve, we can expect to see new best practices and technologies emerge. Organizations that are considering implementing data mesh should stay informed about these developments and carefully evaluate whether this approach is right for them. In the context of cloud databases, data mesh can be a powerful tool for managing and leveraging data in a cloud environment. By distributing data ownership and management across different domains, data mesh can help organizations to take full advantage of the scalability and flexibility of the cloud.
Works cited
- Zhamak Dehghani | Thoughtworks United States, accessed January 10, 2025, https://www.thoughtworks.com/en-us/profiles/z/zhamak-dehghani
- Zhamak Dehghani – The Knowledge Graph Conference, accessed January 10, 2025, https://www.knowledgegraph.tech/speakers/zhamak-dehghani/
- Reflections with Zhamak Dehghani, Founder & Author of Data Mesh, accessed January 10, 2025, https://www.soda.io/podcasts/data-dream-team-zhamak-dehghani-data-mesh-part-one
- Data mesh – Wikipedia, accessed January 10, 2025, https://en.wikipedia.org/wiki/Data_mesh
- The difference between a data mesh and data warehouse – Starburst, accessed January 10, 2025, https://www.starburst.io/blog/data-mesh-vs-data-warehouse/
- Data Mesh Defined: Principles, Architecture, and Benefits – Astera Software, accessed January 10, 2025, https://www.astera.com/type/blog/data-mesh/
- What Is A Data Mesh — And How Not To Mesh It Up – Monte Carlo Data, accessed January 10, 2025, https://www.montecarlodata.com/blog-what-is-a-data-mesh-and-how-not-to-mesh-it-up/
- Exploring Data Mesh: A Paradigm Shift in Data Architecture – KDnuggets, accessed January 10, 2025, https://www.kdnuggets.com/exploring-data-mesh-a-paradigm-shift-in-data-architecture
- What is a Data Mesh? – Data Mesh Architecture Explained – AWS, accessed January 10, 2025, https://aws.amazon.com/what-is/data-mesh/
- An introduction to data mesh – IBM Developer, accessed January 10, 2025, https://developer.ibm.com/articles/data-mesh-a-peek-into-this-new-paradigm/
- Data Mesh Market Primer – K2view, accessed January 10, 2025, https://www.k2view.com/what-is-data-mesh/
- The Top 3 Data Mesh Challenges — and How to Solve Them – Ascend.io, accessed January 10, 2025, https://www.ascend.io/blog/the-top-three-data-mesh-challenges-and-how-to-solve-them/
- What is a Data Mesh? Architecture & Best Practices – Qlik, accessed January 10, 2025, https://www.qlik.com/us/data-management/data-mesh
- What Is a Data Mesh? | IBM, accessed January 10, 2025, https://www.ibm.com/think/topics/data-mesh
- www.starburst.io, accessed January 10, 2025, https://www.starburst.io/blog/data-mesh-vs-data-lake/#:~:text=Data%20lakes%20are%20centralized%20repositories,more%20collaborative%20and%20scalable%20approach.
- Data Mesh vs. Data Lake: The Ultimate Guide to Modern Data Architectures | by Grow.com, accessed January 10, 2025, https://medium.com/@grow.com/data-mesh-vs-data-lake-the-ultimate-guide-to-modern-data-architectures-f17ce4182342
- atlan.com, accessed January 10, 2025, https://atlan.com/data-mesh-vs-data-warehouse/#:~:text=Data%20warehouse%2C%20a%20long%2Dstanding,manage%20their%20data%20as%20products.
- Data Mesh Vs Data Warehouse: 3 Key Differences – Monte Carlo Data, accessed January 10, 2025, https://www.montecarlodata.com/blog-data-mesh-vs-data-warehouse/