Highlightmartinfowler.com

Data mesh scales ownership and change, not only data volume

Zhamak Dehghani's essay on four data mesh principles and the logical architecture each one implies.

Dehghani argues that past technology fixed volume scale but not change, source proliferation, use-case diversity, or response speed. Data mesh answers with four principles: domain-oriented decentralized ownership, data as a product, self-serve data infrastructure as a platform, and federated computational governance. Each principle implies a corresponding logical architecture and organizational structure.

Based on: Data Mesh Principles and Logical Architecture · martinfowler.com

Highlightdbt Labs

manifest.json is the full parsed map of a dbt project's resources

Artifact reference for manifest.json: contents, producers, version mapping, and uses in docs and state comparison.

Commands that parse a dbt project write manifest.json under target/, except deps, clean, debug, and init. The file holds a full representation of resources and properties even when only some nodes run. dbt uses it for the docs site and state comparison; community tools use it for project-health checks. Top-level keys include nodes, sources, metrics, exposures, groups, macros, docs, and parent/child maps.

Based on: Manifest JSON file | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Data tests are select queries that must return zero failing rows

Guide to asserting uniqueness, non-null, relationships, and custom logic on dbt models and related resources.

Data tests are assertions about models and other project resources; dbt test reports pass or fail for each. Built-in checks cover non-null, unique, referential correspondence, and accepted values, and any select that returns failing records can become a test. Generic tests are defined with test blocks; zero failing rows means pass. The tests key remains an alias for data_tests.

Based on: Add data tests to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Exposures name the dashboards and apps that depend on your DAG

dbt docs on declaring manual or automatic exposures that link downstream uses to models, sources, and metrics.

Exposures describe downstream uses of a dbt project such as dashboards, applications, or data science pipelines. Defining them lets you run and test feeding resources and publish consumer-facing pages in generated docs. They can be declared in YAML or created automatically for supported integrations and stored in dbt metadata. Required fields include name, type, and owner; depends_on lists refs, sources, and metrics.

Based on: Add Exposures to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt docs combine YAML descriptions with warehouse metadata

Developer Hub page on generating a documentation site from project descriptions and information-schema queries.

dbt can generate project documentation and render it as a website so downstream consumers can discover curated datasets. Docs cover project information such as model code, DAG, and tests, plus warehouse details like column types and table sizes from the information schema. Authors add description keys on models, columns, sources, and related resources before generating the site.

Based on: About documentation | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Consistent dbt structure frees teams to decide on hard problems

dbt Labs guide arguing that files, folders, naming, and patterns reduce decision fatigue in collaborative projects.

The guide says analytics engineering is collaboration at scale under limited decision bandwidth, so projects need consistent norms for folders, names, and patterns. Structure is how transformations are labeled, grouped, and combined. One shared principle is moving data from source-conformed shapes toward business-conformed ones, with consistency valued over any single style.

Based on: How we structure our dbt projects | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt-core 1.5 ships model contracts, groups, and access controls

GitHub release notes for dbt-core v1.5.0 listing breaking changes and features such as contracts, groups, and access.

dbt-core 1.5.0, released 27 April 2023, adds native data type constraints, model contracts for tables, views, and incremental models, group resources, access attributes, and selection by group. It also deprecates log-path and target-path in dbt_project.yml and allows --select and --exclude multiple times. Private models cannot be ref'd across groups.

Based on: Release dbt-core v1.5.0 · dbt-labs/dbt · dbt Labs

Highlightdbt Labs

Cross-project refs treat public models as a stable API, not a package

dbt Labs documentation on packages versus project dependencies that resolve public models through a metadata service.

dbt long supported installing other projects as packages, which pulls in full source code for macros and models. It also supports project dependencies that resolve on-the-fly refs to public models via a metadata service, so downstream teams do not parse or run upstream models. Those models are treated as an API whose maintainer guarantees quality and stability, with Enterprise prerequisites including public access and a production manifest.

Based on: Project dependencies | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt Mesh is a multi-project pattern for governed cross-team data products

2023 docs introducing dbt Mesh: cross-project refs, Catalog, groups, access, versions, and contracts for independent yet aligned teams.

The page says a single dbt project struggles at scale to coordinate stakeholders, and Mesh addresses multi-project dependencies, governance, and workflows. Mesh is a pattern, not one product: Enterprise cross-project ref, Catalog lineage, governance, groups, access, model versions, and contracts. It recommends treating models as stable APIs when coordinating across teams and outlines when multi-project architecture becomes appropriate.

Based on: About dbt Mesh | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt constraints validate table data only when contracts are enforced

Reference on platform constraints in dbt: validation on write, contract prerequisite, and uneven enforcement across warehouses.

Constraints are platform features that validate data as tables are created or updated; failed validation rolls back the operation. In dbt they apply only to table and incremental models that declare and enforce a contract with explicit column data types. Enforcement varies: some constraints block builds, some are metadata-only, and some platforms cannot define certain types at all.

Based on: constraints | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt model access is group-scoped visibility, not user permissions

Docs distinguishing model access from user access, and explaining groups that mark models private or public for ref boundaries.

The page separates “model access” from dbt user permissions: developers in a project see private models; others may depend only on public ones. Groups give models a shared owner and make interface boundaries explicit. Private access restricts which models other groups may ref; public models are the intended dependency surface, including for future multi-project collaboration.

Based on: Model access | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt model contracts enforce YAML column names, types, and constraints

Reference docs for dbt’s contract config: enforced schema match for supported SQL materializations, with type aliasing notes.

When a contract is enforced, dbt requires the model’s returned dataset to match YAML-defined column names, data types, and supported constraints. The goal is predictable columns for downstream users inside and outside dbt, because even a boolean-to-integer type shift can break queries. Support is limited to certain SQL materializations and platforms; Python models, ephemeral models, and several other resource types are excluded.

Based on: contract | dbt Developer Hub · dbt Labs