HighlightAirbyte

Airbyte Protocol standardizes source and destination actors for ELT pipelines

Airbyte docs defining the protocol: actors, catalog/stream/field primitives, and STDIO JSON vs Socket Protobuf data channels.

The Airbyte Protocol specifies standard components and interactions that declare an ELT pipeline. Sources and destinations are actors with standard interfaces; data is described via catalog, configured catalog, stream, configured stream, and field. Two channel modes exist: STDIO with serialized JSON messages, and Socket mode with Protocol Buffers over Unix domain sockets.

Based on: Airbyte Protocol | Airbyte Docs · Airbyte

HighlightLinkML / Berkeley Bioinformatics Open Source Projects

LinkML authors YAML schemas that compile and validate across JSON, RDF, and TSV

LinkML documentation: a Linked Data Modeling Language for YAML schemas, multi-format validation, and generators into other frameworks.

LinkML is a flexible modeling language for authoring YAML schemas that describe data structure. It is also a framework for working with and validating data in formats such as JSON, RDF, and TSV, with generators that compile schemas to other frameworks. It is Apache-2.0 licensed and community-driven.

Based on: linkml documentation · LinkML / Berkeley Bioinformatics Open Source Projects

HighlightOracle MySQL

MySQL CHECK constraints encode boolean row rules with optional enforcement

MySQL 8.4 reference for CREATE TABLE CHECK constraints: naming, boolean expr, ENFORCED/NOT ENFORCED, table vs column form.

MySQL 8.4 permits core table and column CHECK constraints for all storage engines. A CHECK (expr) must evaluate to TRUE or UNKNOWN per row; FALSE is a violation. Constraints may be ENFORCED or NOT ENFORCED, named with an optional symbol, and declared as table- or column-level.

Based on: MySQL :: MySQL 8.4 Reference Manual :: 15.1.20.6 CHECK Constraints · Oracle MySQL

HighlightPostgreSQL Global Development Group

PostgreSQL COMMENT attaches one replaceable note to almost any database object

Official PostgreSQL docs for the COMMENT command: store, replace, or clear a single comment string on tables, columns, functions, and many other objects.

COMMENT ON stores one comment string per database object and replaces any existing comment when reissued. Specifying NULL or an empty string removes the comment. The synopsis lists object kinds from tables and columns through functions, operators, schemas, and more.

Based on: COMMENT · PostgreSQL Global Development Group

HighlightApache Software Foundation

Iceberg evolves schema and partitions in place without rewriting data

Apache Iceberg docs on metadata-only schema changes and partition evolution that leave existing files intact.

Iceberg supports in-place table evolution for schemas—including nested structures—and for partition layouts as data volume changes, without rewriting data or migrating tables. Schema operations include add, drop, rename, widen types, and reorder; updates are metadata changes tracked by unique column IDs with correctness guarantees. Partition specs can change while old data keeps its prior layout.

Based on: Evolution - Apache Iceberg™ · Apache Software Foundation

Highlightdbt Labs

dbt Discovery API turns run metadata into queryable project state

dbt Labs docs on the Discovery API for querying models, sources, nodes, and run results from dbt projects.

Each dbt run stores metadata about models, sources, other nodes, and execution results. The Discovery API lets you query that metadata to understand the DAG and the data it produces, and to build monitoring, alerting, lineage exploration, and automated reporting. Access is via ad hoc queries, custom apps, partner integrations, and dbt features such as model timing and data health tiles, at environment or job scope.

Based on: About the Discovery API | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Groups bind DAG nodes to a named owner and private-access boundary

dbt documentation on declaring groups in YAML to organize nodes and restrict access to private models.

A group is a named collection of nodes in a dbt DAG with a required owner. Groups support intentional collaboration by restricting access to private models. Members may include models, tests, seeds, snapshots, analyses, and metrics, but not sources or exposures, and each node belongs to only one group. Groups are declared under a groups key; name and owner are required, with optional description and meta in later versions.

Based on: Add groups to your DAG | dbt Developer Hub · dbt Labs

Highlightmartinfowler.com

Data mesh scales ownership and change, not only data volume

Zhamak Dehghani's essay on four data mesh principles and the logical architecture each one implies.

Dehghani argues that past technology fixed volume scale but not change, source proliferation, use-case diversity, or response speed. Data mesh answers with four principles: domain-oriented decentralized ownership, data as a product, self-serve data infrastructure as a platform, and federated computational governance. Each principle implies a corresponding logical architecture and organizational structure.

Based on: Data Mesh Principles and Logical Architecture · martinfowler.com

Highlightdbt Labs

manifest.json is the full parsed map of a dbt project's resources

Artifact reference for manifest.json: contents, producers, version mapping, and uses in docs and state comparison.

Commands that parse a dbt project write manifest.json under target/, except deps, clean, debug, and init. The file holds a full representation of resources and properties even when only some nodes run. dbt uses it for the docs site and state comparison; community tools use it for project-health checks. Top-level keys include nodes, sources, metrics, exposures, groups, macros, docs, and parent/child maps.

Based on: Manifest JSON file | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Data tests are select queries that must return zero failing rows

Guide to asserting uniqueness, non-null, relationships, and custom logic on dbt models and related resources.

Data tests are assertions about models and other project resources; dbt test reports pass or fail for each. Built-in checks cover non-null, unique, referential correspondence, and accepted values, and any select that returns failing records can become a test. Generic tests are defined with test blocks; zero failing rows means pass. The tests key remains an alias for data_tests.

Based on: Add data tests to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

Exposures name the dashboards and apps that depend on your DAG

dbt docs on declaring manual or automatic exposures that link downstream uses to models, sources, and metrics.

Exposures describe downstream uses of a dbt project such as dashboards, applications, or data science pipelines. Defining them lets you run and test feeding resources and publish consumer-facing pages in generated docs. They can be declared in YAML or created automatically for supported integrations and stored in dbt metadata. Required fields include name, type, and owner; depends_on lists refs, sources, and metrics.

Based on: Add Exposures to your DAG | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt docs combine YAML descriptions with warehouse metadata

Developer Hub page on generating a documentation site from project descriptions and information-schema queries.

dbt can generate project documentation and render it as a website so downstream consumers can discover curated datasets. Docs cover project information such as model code, DAG, and tests, plus warehouse details like column types and table sizes from the information schema. Authors add description keys on models, columns, sources, and related resources before generating the site.

Based on: About documentation | dbt Developer Hub · dbt Labs