Article
An introduction to data mesh
Peek into this new paradigmThe term data mesh has been the recipient of a lot of attention in the world of data over the past year and a half. Interestingly, when someone mentions data, our brains connect it to analytical data. Common constructs that come to mind are data warehouses, data lakes, ETL (Extract Transform, and Load), business intelligence, data science etc.
In this article, we stick to the pervasive use and mean analytical data when we mention "data" or "data estate." If Google Trends can be treated as an indicative benchmark, the interest in the data community has been steadily rising in this term. Coined by Thoughtworks, this ideology has been in use for three years.
In this article, we will:
- Look at the data journey of various enterprises at large
- Explain the challenges most enterprises face today with their data estates
- Unpack data mesh and understand how it claims to make things better
- Discuss some headwinds and tailwinds
The data journey
Businesses have been utilizing data for analytical purposes for decades. During the early years of IT, business reporting was done directly from the operational systems that supported the core business of an enterprise. Analytical reporting was generally considered a second-class citizen and was done during non-business hours.
Data warehouse, a structured data source
As the need for business reporting gained importance, data warehouses were created. Enterprises built data warehouses to separate the concerns of two distinct user personas: the operational system users and the analytical users. As data warehouses matured, enterprises attempted to move their operational data into central data warehouses which were purposely built for business intelligence and advanced data analytics. Multiple ETL pipelines fed the data warehouse sophisticated data models under an overarching philosophy of storing the maximum amount of information in the smallest storage footprint. The infrastructure that hosted the data warehouse was very expensive and data analytics was now a critical skill.
Data lake, a structured and unstructured data source
Let us fast forward a few years. Two major changes, the emergence of cloud and its adoption by enterprises and the development of massively parallel processing (MPP) engines, created the need for a new construct, the data lake. The overarching philosophy of data lake was to store data without worrying about its shape and size and transform it as per the use cases to promote artificial intelligence and machine learning (AI/ML) adoption. Most enterprises struggled with both these architectural patterns. Some attempted to build a data warehouse on a data lake while others tried to satisfy data lake use cases on the existing data warehouse. This led to very complex data estates.
The polyglot architecture
Move ahead a few years and two additional factors further influenced the data estates.
- With advancement in technology and a cultural shift towards open source and collaborative software development, data sources proliferated. Now, data came in various shapes, sizes and speed and from various sources, all landing into the central data lake.
- The data's consumers wanted the latest data, accessible in more than one format, and a choice in analytical tools to carry out their tasks.
Sourcing the latest data was handled by Lambda and Kappa pipelines feeding the data lake or the data warehouse. Data was made available in multiple formats e.g., relational tables, REST APIs, graph databases, document databases etc. to users.
Enterprises quickly realized that a polyglot architecture was needed to acquire, organize, and serve data in such a complex scenario. Other paradigms such as data fabric and data lakehouse came into existence with their own merits and their unique solutions to improve the data estate. However, data estate is still not a fully solved puzzle and has only grown more complex.
Challenges facing enterprises today
Enterprises have come a long way from where they started decades ago, in their data journey. However, there are still inhibitors that prevent the enterprises from leveraging the full value of data. We have noted the major ones below:
- Discoverability -- Few enterprises have been able to mature their data estate to a level that they can set up a data marketplace where data consumers can search, understand, and make informed decision about datasets that they want to use.
- Ownership -- There are often challenges in establishing ownership of a dataset. Who owns the dataset and who can certify that a dataset is trustworthy? Most often, the IT teams that own the data platform end up becoming the custodian of the data even though they may not understand it completely.
- Productivity -- Data analyst and business analysts spend 30% to 40% of their time searching for the right dataset. Data engineers spend a major portion of their time determining how to join disparate sources to create a semantically uniform dataset.
- Agility -- In large enterprises, change is constant. Data estates are not able to keep pace with these fast changes and have become an inhibitor to enterprise agility. A new report generation take weeks, which is a lot in the current fast changing world.
- Skills -- Almost the entire data workforce needs to have specialized skills. Maintaining it becomes expensive and missing skills become a bottleneck quite often.
- Trustworthiness -- Quality, observability, and traceability are still some facets that needs robust implementation. Can I trust the data? Am I using the latest file; is it complete? Is it coming from the right source? These are difficult questions, and you hardly get convincing answers readily.
- Self Service -- Self-service capabilities like platforms, datasets, and tools are either not available or even if available are not enough in many enterprises.
Data mesh views these challenges from a completely different lens and attempts to, if not solve, reduce the severity of these problems. Let's dig into what data mesh is.
Data mesh, a new paradigm
The ideology of data mesh suggests that the current data estate challenges faced by enterprises cannot be solved by throwing more technology at them.
The solution lies in reorganizing the three major players in the enterprise: the people, the processes, and the tools.
Thoughtworks, calls data mesh a "socio-technical" paradigm. How does a large enterprise, having complex interconnected systems, stay agile and yet derive maximum value out of its data estate?
Data mesh proposes that enterprises need to start looking at their data estate from a unique perspective encapsulated in the four foundational principles. The principles are elaborated in the subsequent sections to provide insight into the data mesh construct. These principles are inter-related, and they should be understood in the same sequence as outlined.
The four data mesh principles:
- Decentralized ownership of data
- Data as a product
- Self-serve platform
- Federated computational governance
Decentralized ownership of data
This principle is primarily focused on the people and advocates for bringing the operational and analytical worlds closer together. It advocates for the remodeling of the monolith analytical organization by decentralizing and realigning the ownership of the analytical data from a central team to domains.
Most enterprises are already organized into domains and sub-domains to better address the ever-changing needs of the market. A domain can be a business area within an enterprise having a logical boundary. Domains can be lines of businesses, geography of operations, product lines, or even smaller business areas that provides a bounded context. When the operational world transitioned into domain-driven service mesh, it aligned itself to these domains. The analytical world is still alienated and this principle advocates extending the domain concepts to the analytical world.
What problems does decentralized data ownership attempt to solve?
- Ownership -- Many times, data owners are not known. IT teams supporting the ETL, or the platform becomes the de facto owners of the data. In many cases, central IT teams act as the intermediaries by taking requests from the consumers and passing it to the producers. They are not considered the owners primarily because a) they do not produce the data b) they do not understand it. The solution lies in realigning the ownership of the analytical data to the respective domains as they are the primary producer of the data and can understand it the best.
- Agility -- Enterprises have become slow to respond to market as for any business change to take effect, changes must be made across multiple IT systems. Coordination between the team and non-aligned priorities across teams are impediments to enterprise agility. With the proliferation of the data sources and the ever-growing business use cases, central teams in the analytical world have become bottlenecks. Data mesh builds upon the practices that have already been successfully applied in the operational systems. The paradigm shift from one monolithic central place of truth to domain driven microservices has helped operational systems be more agile. Data mesh indents to extend the same concept to the analytical space.
- Productivity -- Data consumers spend most of their time finding the right owner, establishing the traceability, and interpreting the meaning of data. These time-consuming efforts reduce the overall productivity of the teams in the analytical world. Decentralization brings the operational and the analytical world closer under one sphere of control (i.e., domains, thus establishing clear ownership, traceability and a common, clear lingo thereby improving the turnaround time by the teams).
This principle attempts to solve the problem related to ownership of data by bringing in domain centric accountability but there are still so many pieces of the jigsaw that still needs to be solved. Let us say, domains are identified, and they start owning the data. What next? Should they simply keep it in a secure vault going by the popular saying "data is an asset"? Would it not create multiple silos in the enterprise? How would data consumers find the data they are looking for? What happens to analytical insights that depend on cross domains? The second principle of data mesh answers some of these questions.
Data as a product
This principle is primarily focused around the people and the processes. It calls for a mindset shift in the enterprise, a shift from considering analytical data as an asset that should be stored to considering analytical data as a product that should be served.
The domains should consider analytical data as a first-class product rather than considering it a by-product of their business operations. They should also apply all the aspects of product development to make it valuable, useful, reliable, and customer-focused.
A few broader-level aspects need to be understood while considering the principle of data as a product:
- A data product is an autonomous architectural quantum that forms the fundamental building block of the data mesh.
- A domain can have one or more data products.
- Data products should interoperate; it can consume outputs from other data products and produce its own output. Eventually, multiple data products interacting with one another will form a mesh of data products.
- All the technical plumbing such as sourcing data, data modeling, ETL etc. will be abstracted under the data product. The data product team will have authority to design and implement the technical solution.
- Some data products will be aligned to source systems or operational systems. For example, in retail banking industry, current account savings accounts (CASA) can become a data product. CASA data product can produce output such as real-time account transactions, account balances, monthly expenditure, monthly income etc.
- Some data products will consume output from the source aligned data products along with other data products and generate value added output. Extending the same example, categorization of spending on transactions.
- Some data products will be aligned to the extreme right side of the value chain. An example could be a data product producing data for BI report.
This is a major change from the way enterprises think about analytical data. For bringing these changes certain new roles must be carved out with-in the domains. The most important role is the data product owner. This role is responsible for:
- Creating the vision and the feature roadmap of the data product
- Customer satisfaction and ensuring the data product is used within the enterprise
- Ensuring availability, quality, traceability, and maintaining service levels
What problems does product thinking attempt to solve?
- Trustworthiness -- With ownership of data realigned to domain, the data product owner is accountable for the data product. The data product owner ensures that the quality, traceability, and security of the data product is maintained and reported through appropriate metrics, SLOs etc.
- Discoverability -- Each product is cataloged and advertised on the enterprise data marketplace and is self-explanatory. The documentation clearly explains the usability topics such as interfaces, schema, business schematics, and relationship with other data products and SLOs. This ensures data consumers get complete visibility about the data product and they can take informed decisions regarding its usage.
- Agility -- A data product is an autonomous unit of architecture that has its own independent feature roadmap and release cycles. Data products team do not wait for some central platform team to provision the environment and provide data for their work to begin. There is no wastage of time in establishing authenticity, traceability, or in re-work due to non-alignment of input dataset SLOs with the use case SLOs.
- Productivity -- Data consumer productivity automatically increases when aspects like agility, discoverability and trustworthiness are taken care of.
Although the concept of data product seems to bring in many benefits, it might result in an increase in the overall operating cost due to multiple independent infrastructures and multiple small highly skilled teams. Suboptimal utilization of highly skilled teams will drive the operating cost high. The third principle of data mesh attempts to address some of these challenges.
A self-serve platform
This principle is majorly focused around the people and the tools. It states that enterprises should invest in a central data infrastructure to facilitate data product life cycle. Central infrastructure should be self-serve and support tenancy to facilitate autonomy, at the same time, provide multiple tools out of the box. Self-serve tools need to be very thoughtfully built. With the objective of reducing the overall cognitive load on data product teams, they should bring enough abstraction over low level technical components, facilitating faster development and standardization of data product.
The self-serve platform should do the following:
- Be built and maintained by a smaller, central, highly skilled team -- It should be made available to the data product teams as a service through a subscription model.
- Provide standard purpose-built tools that reduce the cognitive load on the data product teams -- Such tools will lower the need for maintaining a large highly skilled team. The team could be composed largely of generalists with a small set of specialists. Examples of such tools include simple user interfaces to define and register events and event schema, to provision a messaging platform, to provision serverless data streaming pipelines, to design and test the transformation scripts. The idea is to bring abstraction to a level where most of the complex functions are hidden and automated.
- Support multi-tenancy and should be able to onboard tenants -- Tenants in this case are data products. This is similar to what the cloud providers have been doing for a while now.
- Provide standard input and output interfaces -- Standard consumer integration patterns and standard producer integration patterns. The input and output could be in the form of data files, data APIs, data streams, etc.
- Automate service provisioning -- Enterprise level cross-cutting concerns such as cost, security, regulatory support, machine learning feature store, data marketplace, etc. should be supported by the platform out of the box.
- Provide polyglot storage options and processing options -- Examples could be relational storage, key value storage, in-memory data grid, SQL querying engines etc.
- Attempt to converge the technology underlying the operational application and analytical applications -- For example, if operational applications are implemented as microservices and are deployed on Kubernetes, streaming applications can also be deployed on Kubernetes as the orchestrator. Spark-based batch processing can also happen on a Kubernetes based cluster.
- Provide AI/ML toolkit and support MLOps -- Provision environment with popular data science tools such as notebook, libraries such as TensorFlow, XGBoost, Keras, etc. that facilitate model training and MLOPs that eases model deployment.
What problems does the self-serve data platform solve?
- Agility -- Autonomous data product team could directly use self-service platform and does not have to depend upon the central infrastructure team to provide the data and infrastructure resources. This leads to a faster development cycle of data products.
- Cost of Ownership -- From an infrastructure perspective, the cost of ownership reduces because it is still centrally provisioned.
- Skills -- Since the platform abstracts technical complexity, the composition of teams shifts toward more generalists and fewer specialists. This reduces the need for a large highly skilled team. |
These principles, if implemented appropriately, seem to address most of the challenges that the enterprises are currently facing. There is still an area that needs to be considered. Most of the data products need to operate across domains. How to determine cust_id in domain A is same as entity_customer_no in domain B? How should you harmonize data across domains? And thus comes the discussion on governance modeling which takes us to the last principle of data mesh.
Federated computational governance
This principle involves all the three major players in the enterprise, the people, the process, and the tools. Federated computational governance is a major shift in ideology from the conventional central governance implementations. The shift is is related to the following ideology concepts:
- How governance teams are organized
- How infrastructure should support governance
- Systemic approaches that can be leveraged to control loosely coupled federated player
Governance should be divided into local governance and global governance:
- Local governance body is local to a data product and is responsible for defining the local governance policies, frameworks, processes and is accountable for their implementation alongside constant adherence. This implementation is, in a sense, moving away from central governing bodies that used to create policies, validate, and certify adherence on the data estate. In federated governance, aspects like data quality, data modeling, local access policies etc. are managed by the data product owner. This is a significant shift, from implementing large canonical data models to smaller data models purpose build to serve the requirements of the data product. Logical extension of the model is achieved by another data product which takes input from the first data product and builds functionality on top of it to generate value.
- Global governance body is a thin, cross-functional body that has subject matter experts in various specializations in the enterprise such as legal, security, domains, infrastructure, and technology. They formulate policies that are overarching and are required for the data products to interoperate. Some examples could be a) the decision on which data product aligns with which domain b) agreement on global attributes so that data products can join or union data from multiple data products c) legal policies such as GDPR, Sox d) Security policies related to encryption, obfuscation, masking e) data classification policies. Global governance body is responsible for formulating the policies and the local governance body is accountable for implementation and constant adherence.
Data mesh will be in a constant state of change as new data products will be launched and older ones will retire. Data mesh advocates de-centralization of governance and proposes that system thinking should be applied to bring about an equilibrium between data product autonomy and mesh harmony. This is a deeper topic for discussion but simply put, instead of global governance dictating which data products should be launched, indicators such as data product usage statistics, consumer satisfaction, etc. can decide the life of the product. An analogy could be courses at Udemy. We find many courses on the same topic, however, the user statistics decide the longevity of the course. Similarly, operational metrics such as data product downtime, change failure rate decide the stability of the data product and whether it needs to be redesigned.
There is a lot of emphasis on computational in this principle. As much as possible, the governance policies -- local or global -- should be automated, coded in the platform, and executed automatically. Platform SMEs are part of the global governance, and they ensure that standards-as-code, policy-as-code, automated test cases, and automated observability is implemented. Examples include automated documentation framework, automated masking/obfuscations, mandatory test cases, automated observability metrics calculation; reporting standard authentication framework, standard data transfer frameworks, and system audit reports.
All these principles are essential for holistic implementation of data mesh in an enterprise. The degree of implementation can vary but as we have seen, each principle overcomes the shortcomings of the other. The larger the mesh the greater the value generated from the data. One could question though; a larger mesh equals a complex mesh. We agree, however we also believe that to make the external interfaces of a system simple and easy to use, the internals of the system must be overly complex. Look at the human beings and their anatomy!!
Headwinds and tailwinds
Enterprises opting to adopt data mesh should be ready to accept the fact that a true implementation would increase the complexity of the data landscape by an order of magnitude. Since data mesh is a philosophy, it will physicalize differently for different enterprises and likewise the degree and the area of complexity. Nevertheless, there are some common areas which needs deeper thinking and careful handling.
Headwinds -- Organizational realignment
With decentralized ownership, the analytical team needs to be completely reorganized at all levels. Careful consideration must be given to define boundaries, operating model across boundaries, roles and their KRAs, reporting structures and allocation of budgets. The degree of acceptance of change varies from individual to individual and here we are talking about a major realignment. Motivating a large team to adopt to this change could be daunting task. Executive promotions, road shows, awareness drives, incentive programs, and many more, all such techniques will play key role.
Headwinds -- Loosely coupled data models with interoperable interface definitions
Enterprises are used to seeing reference data models, tailored towards an industry, that are large, complex, and monolithic. They are designed and maintained by a team of highly skilled data modelers. Data mesh advocates moving away from such a construct because it is difficult to build and non-agile. It also proposes the adoption of a distributed master architecture which is contrary to the golden source, master data management philosophy. This could mean that customer 360 could be a consortium of multiple data products, instead of being a single analytical application. We think it's a subject for deeper thinking.
Headwinds -- Proliferation of data products and associated pitfalls
The larger the mesh the greater the value. Is it that simple? Data mesh encourages a systemic approach to govern the life cycle of the data products. Systemic selection works well in a lightly regulated marketplace. It is lightly regulated because the regulator does not promote/demote any product neither does it define the offerings of any product. Is such a mechanism suitable within the boundaries of an enterprise? Can it result in proliferation of data products, similar products, duplicate products, low value products? How do we justify the money spent on these data products? Enterprises might adopt a more centralized governance and steadily gravitate towards more federated governance as the data mesh culture matures.
On one hand, there are areas that need deeper thinking to physicalize data mesh. On the other hand, there are aspects that amalgamates very well with few recent trends that enterprises are embracing.
Tailwinds -- Data first
We have named it the data-first approach. Remember that data is synonymous to analytical data. While defining a business process, inspect each process to define analytical measures that are relevant to improve the business process. This is what we call data-first approach.
Enterprises want to be driven more and more by data. They need analytical insights about their business process in the quickest amount of time possible. This enables enterprises to adapt and respond to changes faster. Traditionally, defining insights was always done at a later point in time, much later than the implementation of the business process. This trend is changing, and we clearly understand the rationale.
Defining and building analytical insights about operational processes early on, is forcing the analytical world and the operational work to converge. In our opinion, source-aligned data products are the most appropriate candidates to either produce relevant data elements required for the analytical measures or if possible, produce the analytical measures itself. This is how we feel that data mesh and data first align very well.
Tailwinds -- AI-infused
Enterprises are increasingly embracing AI and including AI as a core component while designing or transforming their business processes. As a result, enterprises are focusing heavily on building and maintaining an ever-expanding feature store. On one hand, these features stores are used for training AI/ML models implemented in their existing business process, on the other hand, they are also used for creating new and innovative use cases for the enterprises. The question to consider is, how do we build such features stored and keep hydrating them? Furthermore, quite a portion of this feature store could be sourced from unstructured data emanating from sources such as chatbot, interactive voice response (IVR) etc. In our opinion, such feature stores can be the outcome of source-aligned data products or even in some cases intermediate data products.
Tailwinds -- Cater to business and IT alike
Business users are now keener than ever to carry out data exploration, exploratory analysis, and self-serve reporting. Enterprises have started focusing on providing self-serve platforms to business users where they get access to production data, fire SQL queries, and are able to create self-service business intelligence (BI) reports. Data mesh proposes a self-serve platform that can be used by the data product team to build data products. While this platform can keep serving the IT, it can also have capability engineered towards serving the business users. By doing this, we are enabling the right set of people who can make a direct impact to the business outcome or an enterprise.
Summary
In this article, we explained the principles of data mesh and how it can solve the current day problems with enterprises data estates. We reviewed a few areas that will need deeper thinking to come up with the right implementation. We also eluded on few recent trends where data mesh becomes a natural fit. We have purposefully restrained ourselves to the what part of data mesh. We intend put together our thoughts about how in the subsequent article. Stay tuned and we hope you have found this useful!
We would like to extend our thanks to Tanmay Ambre for reviewed and providing recommendations on content and structure, and Patty Orben for detailed editing to make the article better organized and more readable.