HighlightApache Software Foundation

Iceberg evolves schema and partitions in place without rewriting data

Apache Iceberg docs on metadata-only schema changes and partition evolution that leave existing files intact.

Iceberg supports in-place table evolution for schemas—including nested structures—and for partition layouts as data volume changes, without rewriting data or migrating tables. Schema operations include add, drop, rename, widen types, and reorder; updates are metadata changes tracked by unique column IDs with correctness guarantees. Partition specs can change while old data keeps its prior layout.

Based on: Evolution - Apache Iceberg™ · Apache Software Foundation

HighlightJSON Schema (IETF-oriented draft)

JSON Schema Validation vocabulary asserts structure, meaning, and UI hints

Internet-Draft specifying the JSON Schema vocabulary for instance validation, document meaning, and UI hints.

This draft specifies a vocabulary for JSON Schema used in JSON instance validation. It describes meanings of JSON documents, hints for user interfaces, and assertions about valid document shape. Authors are Wright, Andrews, and Hutton; the draft was published 16 June 2022 as informational IETF work in progress.

Based on: JSON Schema Validation: A Vocabulary for Structural Validation of JSON · JSON Schema (IETF-oriented draft)

HighlightJSON Schema (IETF-oriented draft)

JSON Schema defines a media type for describing JSON document structure

IETF-oriented Internet-Draft for JSON Schema as application/schema+json, including instance media type notes.

This Internet-Draft defines JSON Schema as the media type application/schema+json, a JSON-based format for describing JSON structure, extraction, and interaction. It also discusses application/schema-instance+json for richer integration than plain application/json. The draft is informational, authored by Wright, Andrews, Hutton, and Dennis, published 16 June 2022.

Based on: JSON Schema: A Media Type for Describing JSON Documents · JSON Schema (IETF-oriented draft)

HighlightGoogle

Proto3 defines typed messages and stable field numbers for shared payloads

Google’s proto3 language guide on .proto syntax, message fields, types, and generating data access classes.

The proto3 Language Guide explains how to structure protocol buffer data with .proto syntax and how generated access classes are produced. A message declares named, typed fields each assigned a field number. The edition or syntax line must be first; if omitted, the compiler assumes proto2.

Based on: Language Guide (proto 3) · Google

Highlightdbt Labs

manifest.json is the full parsed map of a dbt project's resources

Artifact reference for manifest.json: contents, producers, version mapping, and uses in docs and state comparison.

Commands that parse a dbt project write manifest.json under target/, except deps, clean, debug, and init. The file holds a full representation of resources and properties even when only some nodes run. dbt uses it for the docs site and state comparison; community tools use it for project-health checks. Top-level keys include nodes, sources, metrics, exposures, groups, macros, docs, and parent/child maps.

Based on: Manifest JSON file | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Data tests are select queries that must return zero failing rows

Guide to asserting uniqueness, non-null, relationships, and custom logic on dbt models and related resources.

Data tests are assertions about models and other project resources; dbt test reports pass or fail for each. Built-in checks cover non-null, unique, referential correspondence, and accepted values, and any select that returns failing records can become a test. Generic tests are defined with test blocks; zero failing rows means pass. The tests key remains an alias for data_tests.

Based on: Add data tests to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt-core 1.5 ships model contracts, groups, and access controls

GitHub release notes for dbt-core v1.5.0 listing breaking changes and features such as contracts, groups, and access.

dbt-core 1.5.0, released 27 April 2023, adds native data type constraints, model contracts for tables, views, and incremental models, group resources, access attributes, and selection by group. It also deprecates log-path and target-path in dbt_project.yml and allows --select and --exclude multiple times. Private models cannot be ref'd across groups.

Based on: Release dbt-core v1.5.0 · dbt-labs/dbt · dbt Labs

Highlightdbt Labs

dbt constraints validate table data only when contracts are enforced

Reference on platform constraints in dbt: validation on write, contract prerequisite, and uneven enforcement across warehouses.

Constraints are platform features that validate data as tables are created or updated; failed validation rolls back the operation. In dbt they apply only to table and incremental models that declare and enforce a contract with explicit column data types. Enforcement varies: some constraints block builds, some are metadata-only, and some platforms cannot define certain types at all.

Based on: constraints | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt model contracts enforce YAML column names, types, and constraints

Reference docs for dbt’s contract config: enforced schema match for supported SQL materializations, with type aliasing notes.

When a contract is enforced, dbt requires the model’s returned dataset to match YAML-defined column names, data types, and supported constraints. The goal is predictable columns for downstream users inside and outside dbt, because even a boolean-to-integer type shift can break queries. Support is limited to certain SQL materializations and platforms; Python models, ephemeral models, and several other resource types are excluded.

Based on: contract | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt model governance covers access, contracts, versions, and mesh refs

Overview of dbt model governance: public/private access, contracts, versions, namespaces, and Enterprise cross-project dependencies.

The page says model governance controls who can access models, what they contain, how they change, and how they are referenced across projects. Features include public/private access, contracts on column shape, versions for breaking changes, namespaces for ownership, and Enterprise project dependencies via cross-project ref. It also mentions freshness SLAs with State and lag_tolerance, plus caveats about adopting governance too early.

Based on: About model governance | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Model versions treat shared dbt models like versioned APIs

dbt Mesh docs explaining model versioning versus other “version” meanings, and why producers and consumers need graceful change.

The page separates model versions (a Mesh governance feature) from dbt_project.yml and YAML property-file version fields. It compares shared dbt models to APIs where producers must change logic while consumers need stable queries. Model versioning is offered as a deliberate way to handle breaking changes without pretending the tension disappears. It also warns that governance features can harden rollbacks if adopted too early.

Based on: Model versions | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt dimensions add categorical and time attributes to semantic models

dbt documentation for non-aggregatable Semantic Layer dimensions: name, type, optional expr and label.

Dimensions are non-aggregatable attributes in a semantic model—features that categorize data and typically appear in SQL GROUP BY. Each needs a unique name within the model and a type of categorical or time; optional expr and label control column mapping and downstream display.

Based on: Dimensions | dbt Developer Hub · dbt Labs