HighlightAmazon Web Services

AWS Glue Data Catalog indexes location, schema, and metrics

AWS documentation on Glue’s central metadata catalog, crawlers, manual tables, and links to Athena, Lake Formation, EMR, and SageMaker.

The AWS Glue Data Catalog is described as a centralized metadata repository that indexes location, schema, and runtime metrics of data sources in metadata tables. Catalog entries can be filled by crawlers that scan internal and external sources, or defined manually. The catalog feeds ETL and integrates with Athena, Lake Formation, EMR, and SageMaker AI.

Based on: Data discovery and cataloging in AWS Glue - AWS Glue · Amazon Web Services

HighlightGoogle Cloud

Knowledge Catalog: Gemini context graph for agents and governance

Google Cloud overview of Knowledge Catalog (formerly Dataplex Universal Catalog): an AI-powered catalog that builds a context graph to ground agents.

As of April 10, 2026, Dataplex Universal Catalog is renamed Knowledge Catalog while API, client library, CLI, and IAM names stay the same. The product is described as a Gemini-powered catalog that extracts semantics from structured and unstructured data into a dynamic context graph for discovery, policy, and agent grounding to reduce hallucinations.

Based on: Knowledge Catalog overview | Google Cloud Documentation · Google Cloud

HighlightSnowflake

Semantic views put business metrics and entities in the database

Snowflake overview of Semantic Views: schema-level objects that define metrics, entities, and relationships for consistent business meaning and Cortex Agents.

A Semantic View is a schema-level Snowflake object that stores business metrics, entities, and relationships as metadata atop physical data. It supplies consistent definitions across applications, is queryable with SELECT, usable by Cortex Agents, and shareable via listings. The docs argue this layer fixes the gap between business language and opaque column names and stops inconsistent metric calculations across reports.

Based on: Overview of semantic views | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake Information Schema is the per-database data dictionary

Snowflake docs on INFORMATION_SCHEMA views and table functions for database and account-level object metadata.

Snowflake’s Information Schema is a SQL-92 ANSI–based data dictionary implemented as a built-in, read-only INFORMATION_SCHEMA in every database. It exposes views for database objects and account-level objects (roles, warehouses, databases) plus table functions for historical and usage data, mixing ANSI-standard and Snowflake-specific views.

Based on: Snowflake Information Schema | Snowflake Documentation · Snowflake

HighlightAirbyte

Airbyte Protocol standardizes source and destination actors for ELT pipelines

Airbyte docs defining the protocol: actors, catalog/stream/field primitives, and STDIO JSON vs Socket Protobuf data channels.

The Airbyte Protocol specifies standard components and interactions that declare an ELT pipeline. Sources and destinations are actors with standard interfaces; data is described via catalog, configured catalog, stream, configured stream, and field. Two channel modes exist: STDIO with serialized JSON messages, and Socket mode with Protocol Buffers over Unix domain sockets.

Based on: Airbyte Protocol | Airbyte Docs · Airbyte

HighlightPostgreSQL Global Development Group

PostgreSQL COMMENT attaches one replaceable note to almost any database object

Official PostgreSQL docs for the COMMENT command: store, replace, or clear a single comment string on tables, columns, functions, and many other objects.

COMMENT ON stores one comment string per database object and replaces any existing comment when reissued. Specifying NULL or an empty string removes the comment. The synopsis lists object kinds from tables and columns through functions, operators, schemas, and more.

Based on: COMMENT · PostgreSQL Global Development Group

HighlightSnowflake

Snowflake Horizon Catalog pairs discovery with semantic context for agents

Snowflake documentation introducing Horizon Catalog as an interoperable catalog with governance, lineage, and semantic views for AI.

Horizon Catalog is described as an agentic catalog for data inside and outside Snowflake, open across engines, data, and clouds. It targets discoverability, business context for AI, and trust via protection, quality, lineage, and AI governance. Features include Iceberg interoperability, Internal Marketplace, semantic views, and column-level lineage spanning Snowflake, external systems, BI, and OpenLineage.

Based on: Snowflake Horizon Catalog | Snowflake Documentation · Snowflake

Highlightdbt Labs

dbt Discovery API turns run metadata into queryable project state

dbt Labs docs on the Discovery API for querying models, sources, nodes, and run results from dbt projects.

Each dbt run stores metadata about models, sources, other nodes, and execution results. The Discovery API lets you query that metadata to understand the DAG and the data it produces, and to build monitoring, alerting, lineage exploration, and automated reporting. Access is via ad hoc queries, custom apps, partner integrations, and dbt features such as model timing and data health tiles, at environment or job scope.

Based on: About the Discovery API | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt Semantic Layer APIs put MetricFlow metrics behind governed query interfaces

dbt docs on Semantic Layer APIs: define metrics in code with MetricFlow and query governed assets from downstream tools, including via GraphQL.

The page argues that growth in the modern data stack fragments business logic across teams and tools. The Semantic Layer lets you define metrics in code with MetricFlow and dynamically generate and query datasets from dbt-governed metrics and models. It lists use cases such as BI, data quality, governance, discovery, and ML, and notes a GraphQL API for metrics and dimensions.

Based on: Semantic Layer APIs | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt exports materialize saved MetricFlow queries as warehouse tables

dbt docs on exports: run saved Semantic Layer queries via the job scheduler and write results to tables or views.

Exports run saved MetricFlow queries and write output to a table or view in the data platform, using the dbt job scheduler. They give SQL and non–Semantic Layer tools access to metric definitions as ordinary relations; running an export counts toward queried metrics usage, querying the result does not.

Based on: Write queries with exports | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt dimensions add categorical and time attributes to semantic models

dbt documentation for non-aggregatable Semantic Layer dimensions: name, type, optional expr and label.

Dimensions are non-aggregatable attributes in a semantic model—features that categorize data and typically appear in SQL GROUP BY. Each needs a unique name within the model and a type of categorical or time; optional expr and label control column mapping and downstream display.

Based on: Dimensions | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt entities are typed join keys that link semantic models

dbt docs define entities as business concepts used as join keys—primary, unique, foreign, or natural—in the Semantic Layer.

Entities represent real-world concepts such as customers or transactions and act as join keys across semantic models in dbt’s Semantic Layer (v1.12+). Required parameters are name and type; types are primary, unique, foreign, and natural. Names must be unique within a model and may use expr; entities can also be used as dimensions.

Based on: Entities | dbt Developer Hub · dbt Labs