Why centralised data architectures are reaching their limits
Data mesh is a set of principles for designing modern data architectures, formalised by Zhamak Dehghani in 2019. Its starting point is an observation of failure: centralised architectures, the warehouse, then the data lake, break down at organisational scale. A single data team, squeezed between dozens of producing domains and hundreds of consumers, mechanically becomes the bottleneck of the entire company.
The symptoms are well known: months of waiting to onboard a new source, pipelines that break with every upstream change, data whose provenance and reliability nobody can vouch for any more, and business teams that end up rebuilding their own extractions in a corner, recreating the very silos centralisation was supposed to remove. The problem is not technological: it is a problem of ownership, which adding yet another integration tool never solves.
Data mesh inverts the logic: data stays under the responsibility of the domains that produce and understand it, and circulates between them as documented, contract-backed products. It is an organisational model as much as an architectural one, and that is precisely what makes it both powerful and demanding. Adopting the mesh means accepting that the data architecture follows the boundaries of the organisation, and not the other way round.
The four founding principles of data mesh
The four principles form a system: each one compensates for the risks introduced by the others, and applying only one or two of them generally produces more disorder than the centralised architecture you started from. Decentralisation without a platform creates chaos, a platform without governance creates incompatible islands, and governance without domain ownership recreates the central bottleneck under another name.
Decentralised domain ownership
Each business domain, orders, customers, logistics, owns and operates its analytical data, just as it already owns its applications. The team that knows the data best takes responsibility for its quality, its documentation and its evolution, instead of throwing it over the wall to a central team that discovers it without context.
Data as a product
A data product is managed like a software product: it has consumers it must satisfy, a stable interface, documentation, quality and freshness guarantees, and an identified owner. Discoverable, addressable, trustworthy and interoperable: these four qualities turn data from an operational by-product into an asset that other teams can consume with confidence. The satisfaction of internal consumers becomes a steering indicator in its own right, on a par with a service's availability.
The self-serve platform
A platform team provides the shared infrastructure, storage, pipelines, catalogues, access control, as self-service, so that every domain can publish and consume data products without deep infrastructure expertise. Without this platform, decentralisation multiplies costs by the number of domains; with it, the marginal cost of a new data product collapses. The platform team measures itself on the autonomy it gives the domains, not on the tickets it processes.
Federated computational governance
The cross-cutting rules, shared identifiers, classification, security, compliance, are decided collectively by domain representatives, then encoded into the platform and enforced automatically. "Computational" is the key word: governance executes in the pipelines and the catalogues, it does not live in documents nobody reads.
Data contracts and the open lakehouse: the tooling that made the mesh operational
Long a conceptual model, data mesh has been industrialised thanks to two building blocks that have become standard. The data contract, first: a formal, versioned agreement between a producer and its consumers that fixes the schema, the semantics, and the quality and freshness thresholds of a data product. Verified automatically in the pipelines, it turns implicit promises into testable guarantees, a contract breach is detected at publication time, not three weeks later in a wrong dashboard. Modern transformation tools embed these tests natively, which makes the contract a living artefact rather than a statement of intent.
Open formats and interoperable catalogues
Open table formats, Apache Iceberg first among them, whose v3 specification is now supported by the major engines and clouds, decouple storage from compute engines: the same data product can be consumed from several engines without copies or vendor lock-in. Catalogues and automated lineage make products discoverable and traceable across the organisation, the precondition for genuine self-service. Interoperability is no longer a promise on a slide: it is guaranteed by the format specification itself.
The mesh does not condemn the data lake, however: the open lakehouse is often its technical substrate, with each domain operating its own governed tables on it. The difference is organisational: ownership and responsibility are distributed, even when the infrastructure is shared.

Federated governance in practice: GDPR and AI Act compliance without a central bottleneck
Governance is the hardest principle, and the most profitable one. By keeping data in its source systems and making explicit who is responsible for it, data mesh makes it easier to comply with GDPR and sector regulations: responsibilities are clear, exposure is limited, and classification is applied as close to the source as possible. Auditors get a single, consistent answer to the question of who owns which data, often for the first time.
Encoding the rules in the platform rather than in documents
Policies are encoded into the platform: automatic masking of personal data according to classification, access rights propagated along the lineage, tooled purging for the right to erasure. The AI Act raises the stakes: traceability of the data feeding high-risk AI systems becomes a regulatory obligation, and the native lineage of data products is the most direct way to demonstrate it.
In practice, federated governance lives in a lightweight body, representatives of each domain, the platform and compliance, that arbitrates shared standards and settles edge cases. Its effectiveness is measured by what it automates: every rule that moves from a document into a pipeline is one meeting fewer and one compliance incident avoided.

Data mesh and AI: data products for models and agents
The rise of generative AI has given data mesh a second wind. Models, RAG systems and agents are only worth as much as the data you feed them: an agent reasoning over stale or inconsistent data produces wrong answers with complete confidence. Data products, documented, contract-backed, with freshness guarantees, are exactly the interface these systems need. Teams that had invested in the mesh are finding that their data products feed AI use cases without friction, while their competitors spend months hardening ad hoc extractions.
From data product to agent-consumable product
Advanced organisations expose their data products to AI systems the same way they expose them to analysts: through governed interfaces, with semantic metadata and access control. The documentation and explicit semantics of the products, designed for humans, now also serve as context for models, a well-kept catalogue becomes a map that agents know how to read. The same governed interfaces also keep sensitive data out of prompts by construction, rather than by policy reminders. Conversely, plugging an LLM into an ungoverned data lake amounts to industrialising the production of errors.
Field experience: who adopts data mesh, and on what conditions
The approach is not theoretical: Netflix, Zalando, Intuit and VistaPrint adopted it long ago to manage the growing complexity of their data landscapes, and many European banking, industrial and retail groups have followed. The feedback converges: the value comes from shrinking the delay between a business question and its answer, and from the gradual disappearance of shadow pipelines. That delay drops from weeks to a few days, sometimes a few hours, once the products cover the core domains.
The failures converge just as much. Data mesh is not a product you install: without clearly delimited domains, without teams ready to take responsibility for their data and without real investment in the platform, decentralisation degenerates into fragmentation. The organisation's data maturity is the real prerequisite, a mesh is earned, not decreed. The organisations that succeed share one trait: they treated adoption as a product and organisational transformation, with measurable milestones, not as an infrastructure project.

The Adservio approach: aligning domains, platform and governance
At Adservio, we treat data mesh as a discipline of alignment between business and technology: delimiting the domains, defining the first data products and their contracts, standing up the self-serve platform and the federated governance, rather than an isolated technology migration that would change the tools without changing the responsibilities. The initial diagnosis covers the organisation, who owns what, who consumes what, where the frictions are, as much as the existing architecture.
We favour incremental adoption: two or three pilot domains, high-value data products consumed by real use cases, analytics or AI, then a generalisation driven by proof rather than by mandate. Our conviction: data is all the more valuable when it stays close to those who produce and understand it. We equip that proximity with governance and observability, and transfer control to your teams so they operate their data products autonomously. Skills transfer is planned from the first sprint, not bolted on at the end of the engagement.
STAY POSTED
Get our next analyses and field notes straight to your inbox.




