Data mesh: domain-owned data products instead of a central lake monolith
How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh
Zhamak Dehghani’s 2019 essay on moving from monolithic data lakes to a distributed data mesh.
Based on
How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh
Enterprises invest in next-generation data lakes to democratize data and automate decisions, but lake-style platforms share failure modes at scale. Dehghani calls for leaving the centralized lake or warehouse paradigm for a distributed one: domains as a first-class concern, platform thinking for self-serve data infrastructure, and data as a product. Failure modes named include centralized monoliths, coupled pipeline decomposition, and siloed hyper-specialized ownership. Domain data should be discoverable, addressable, trustworthy, self-describing in semantics and syntax, interoperable under global standards, and secured under global access control, owned by cross-functional domain teams.
Data platform and governance teams stuck scaling a single lake get an architectural alternative: push ownership to domains while keeping global standards for interop and security.
Self-describing, interoperable data products are shared meaning across domains—contracts and semantics that let distributed teams consume each other’s data without a single central model.
Put this to work on CoreModels
Related connectors and recipes
Take the next step
Try CoreModels, talk with our team, or explore more resources.