HighlightGoogle

Proto3 defines typed messages and stable field numbers for shared payloads

Google’s proto3 language guide on .proto syntax, message fields, types, and generating data access classes.

The proto3 Language Guide explains how to structure protocol buffer data with .proto syntax and how generated access classes are produced. A message declares named, typed fields each assigned a field number. The edition or syntax line must be first; if omitted, the compiler assumes proto2.

Based on: Language Guide (proto 3) · Google

HighlightMicrosoft

OneLake catalog centralizes find, govern, and secure for Fabric items

Overview of Microsoft Fabric’s OneLake catalog for discovering items, reviewing governance posture, and managing security roles.

OneLake catalog is a centralized Fabric place to find, explore, and use items and to govern owned data. It is reachable from the Fabric navigation pane and embedded in Teams, Excel, and Copilot Studio; metadata can also be searched via the Catalog Search REST API. Explore, Govern, and Secure tabs cover browsing with filters, governance posture insights, and unified workspace and OneLake security role management.

Based on: OneLake catalog overview - Microsoft Fabric · Microsoft

HighlightSnowflake

Snowflake Horizon Catalog pairs discovery with semantic context for agents

Snowflake documentation introducing Horizon Catalog as an interoperable catalog with governance, lineage, and semantic views for AI.

Horizon Catalog is described as an agentic catalog for data inside and outside Snowflake, open across engines, data, and clouds. It targets discoverability, business context for AI, and trust via protection, quality, lineage, and AI governance. Features include Iceberg interoperability, Internal Marketplace, semantic views, and column-level lineage spanning Snowflake, external systems, BI, and OpenLineage.

Based on: Snowflake Horizon Catalog | Snowflake Documentation · Snowflake

Highlightdbt Labs

dbt Discovery API turns run metadata into queryable project state

dbt Labs docs on the Discovery API for querying models, sources, nodes, and run results from dbt projects.

Each dbt run stores metadata about models, sources, other nodes, and execution results. The Discovery API lets you query that metadata to understand the DAG and the data it produces, and to build monitoring, alerting, lineage exploration, and automated reporting. Access is via ad hoc queries, custom apps, partner integrations, and dbt features such as model timing and data health tiles, at environment or job scope.

Based on: About the Discovery API | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Groups bind DAG nodes to a named owner and private-access boundary

dbt documentation on declaring groups in YAML to organize nodes and restrict access to private models.

A group is a named collection of nodes in a dbt DAG with a required owner. Groups support intentional collaboration by restricting access to private models. Members may include models, tests, seeds, snapshots, analyses, and metrics, but not sources or exposures, and each node belongs to only one group. Groups are declared under a groups key; name and owner are required, with optional description and meta in later versions.

Based on: Add groups to your DAG | dbt Developer Hub · dbt Labs

Highlightmartinfowler.com

Data mesh scales ownership and change, not only data volume

Zhamak Dehghani's essay on four data mesh principles and the logical architecture each one implies.

Dehghani argues that past technology fixed volume scale but not change, source proliferation, use-case diversity, or response speed. Data mesh answers with four principles: domain-oriented decentralized ownership, data as a product, self-serve data infrastructure as a platform, and federated computational governance. Each principle implies a corresponding logical architecture and organizational structure.

Based on: Data Mesh Principles and Logical Architecture · martinfowler.com

Highlightdbt Labs

manifest.json is the full parsed map of a dbt project's resources

Artifact reference for manifest.json: contents, producers, version mapping, and uses in docs and state comparison.

Commands that parse a dbt project write manifest.json under target/, except deps, clean, debug, and init. The file holds a full representation of resources and properties even when only some nodes run. dbt uses it for the docs site and state comparison; community tools use it for project-health checks. Top-level keys include nodes, sources, metrics, exposures, groups, macros, docs, and parent/child maps.

Based on: Manifest JSON file | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Data tests are select queries that must return zero failing rows

Guide to asserting uniqueness, non-null, relationships, and custom logic on dbt models and related resources.

Data tests are assertions about models and other project resources; dbt test reports pass or fail for each. Built-in checks cover non-null, unique, referential correspondence, and accepted values, and any select that returns failing records can become a test. Generic tests are defined with test blocks; zero failing rows means pass. The tests key remains an alias for data_tests.

Based on: Add data tests to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Exposures name the dashboards and apps that depend on your DAG

dbt docs on declaring manual or automatic exposures that link downstream uses to models, sources, and metrics.

Exposures describe downstream uses of a dbt project such as dashboards, applications, or data science pipelines. Defining them lets you run and test feeding resources and publish consumer-facing pages in generated docs. They can be declared in YAML or created automatically for supported integrations and stored in dbt metadata. Required fields include name, type, and owner; depends_on lists refs, sources, and metrics.

Based on: Add Exposures to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt docs combine YAML descriptions with warehouse metadata

Developer Hub page on generating a documentation site from project descriptions and information-schema queries.

dbt can generate project documentation and render it as a website so downstream consumers can discover curated datasets. Docs cover project information such as model code, DAG, and tests, plus warehouse details like column types and table sizes from the information schema. Authors add description keys on models, columns, sources, and related resources before generating the site.

Based on: About documentation | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Consistent dbt structure frees teams to decide on hard problems

dbt Labs guide arguing that files, folders, naming, and patterns reduce decision fatigue in collaborative projects.

The guide says analytics engineering is collaboration at scale under limited decision bandwidth, so projects need consistent norms for folders, names, and patterns. Structure is how transformations are labeled, grouped, and combined. One shared principle is moving data from source-conformed shapes toward business-conformed ones, with consistency valued over any single style.

Based on: How we structure our dbt projects | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt-core 1.5 ships model contracts, groups, and access controls

GitHub release notes for dbt-core v1.5.0 listing breaking changes and features such as contracts, groups, and access.

dbt-core 1.5.0, released 27 April 2023, adds native data type constraints, model contracts for tables, views, and incremental models, group resources, access attributes, and selection by group. It also deprecates log-path and target-path in dbt_project.yml and allows --select and --exclude multiple times. Private models cannot be ref'd across groups.

Based on: Release dbt-core v1.5.0 · dbt-labs/dbt · dbt Labs