HighlightAmazon Web Services

AWS Glue Data Catalog indexes location, schema, and metrics

AWS documentation on Glue’s central metadata catalog, crawlers, manual tables, and links to Athena, Lake Formation, EMR, and SageMaker.

The AWS Glue Data Catalog is described as a centralized metadata repository that indexes location, schema, and runtime metrics of data sources in metadata tables. Catalog entries can be filled by crawlers that scan internal and external sources, or defined manually. The catalog feeds ETL and integrates with Athena, Lake Formation, EMR, and SageMaker AI.

Based on: Data discovery and cataloging in AWS Glue - AWS Glue · Amazon Web Services

HighlightMicrosoft Learn

Synapse dedicated pool: pick smallest types to shorten rows

Microsoft Learn recommendations for Synapse SQL Dedicated Pool table data types, row length, PolyBase limits, and finding unsupported types when migrating.

The article gives recommendations for defining table data types in Synapse SQL Dedicated Pool, which supports commonly used types listed in CREATE TABLE. Guidance centers on minimizing row length for performance: smallest workable types, tight VARCHAR lengths, VARCHAR over NVARCHAR when possible, capped lengths instead of MAX, and integer types instead of zero-scale decimals. PolyBase external loads cannot exceed 1 MB per row.

Based on: Table data types in Synapse SQL - Azure Synapse Analytics · Microsoft Learn

HighlightMicrosoft Learn

Fabric lakehouse schemas group tables for access and four-part SQL

Microsoft Fabric docs on lakehouse schemas for domain grouping, schema-level access, cross-workspace queries, and schema shortcuts to external Delta.

Lakehouse schemas in Microsoft Fabric group tables into named collections such as sales or hr. Schema-enabled lakehouses support domain browsing, schema-level access with row- and column-level security, four-part workspace.lakehouse.schema.table queries, schema shortcuts to other lakehouses or ADLS Gen2, and features such as materialized lake views. Schema names allow only letters, numbers, and underscores; dbo is the default under Tables.

Based on: Lakehouse schemas - Microsoft Fabric · Microsoft Learn

HighlightDatabricks

Delta table schema evolution: add, reorder, rename, widen types

Databricks guide to explicit and implicit table schema changes, including DDL column operations and conflicts with concurrent writes and streams.

Databricks tables support schema evolution: adding columns at arbitrary positions, reordering, renaming, and type widening. Changes can be made with DDL such as ALTER TABLE or implicitly via DML. Schema updates conflict with concurrent writes, and updating a schema terminates streams reading the table until they are restarted.

Based on: Update table schemas with schema evolution | Databricks on AWS · Databricks

HighlightDatabricks

Databricks Delta constraints: enforced checks vs informational keys

Databricks documentation of NOT NULL and CHECK enforced constraints versus informational primary, foreign, and unique keys on Delta Lake tables.

Databricks supports enforced constraints that reject violating writes and informational primary key, foreign key, and unique constraints that declare relationships without enforcement. All require Delta Lake. Enforced types are NOT NULL and CHECK; adding constraints may raise the table writer protocol and affect external Delta clients.

Based on: Constraints on Databricks | Databricks on AWS · Databricks

HighlightGoogle Cloud

BigQuery column security via policy and governance tags

Google Cloud intro to BigQuery column-level access control using policy tags or data governance tags, IAM checks at query time, and optional masking.

BigQuery restricts sensitive columns with policy tags from Data Catalog or data governance tags from Resource Manager. Policies are evaluated at query time; optional dynamic masking can replace values with null, default, or hashed content. The workflow is taxonomy and tags, schema annotations on columns, enforce access on the taxonomy, then IAM on each tag, in addition to dataset ACLs.

Based on: Introduction to column-level access control | BigQuery | Google Cloud Documentation · Google Cloud

HighlightSnowflake

Snowflake tags as schema-level key-value labels for governance

Snowflake docs on object tags: schema-level key-value pairs assignable across object types, with inheritance, propagation, and queryable governance use.

A Snowflake tag is a schema-level object stored as a string key-value pair and assigned to other objects. Objects may carry multiple tags; one tag may apply to different object types; values may be shared or unique per assignment. Tags support inheritance down the securable hierarchy, optional automatic propagation, replication of assignments, and centralized or decentralized administration for auditing and reporting.

Based on: Introduction to object tagging | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake Information Schema is the per-database data dictionary

Snowflake docs on INFORMATION_SCHEMA views and table functions for database and account-level object metadata.

Snowflake’s Information Schema is a SQL-92 ANSI–based data dictionary implemented as a built-in, read-only INFORMATION_SCHEMA in every database. It exposes views for database objects and account-level objects (roles, warehouses, databases) plus table functions for historical and usage data, mixing ANSI-standard and Snowflake-specific views.

Based on: Snowflake Information Schema | Snowflake Documentation · Snowflake

HighlightAstronomer

Astronomer on OpenLineage and Airflow for pipeline lineage

Astronomer Learn guide covering data lineage concepts and why integrating lineage with Apache Airflow matters.

The guide defines data lineage as tracking and visualizing data from origin through downstream consumption, and lists uses such as understanding sources, troubleshooting failures, managing PII, and regulatory compliance. It positions Airflow as a central orchestrator for integrating lineage and introduces how lineage works with Airflow, with pointers to Astro Observe and related webinars.

Based on: Integrate OpenLineage and Airflow - Astronomer · Astronomer

HighlightAstronomer

Astronomer guide: schedule Airflow DAGs on asset updates

Astronomer Learn guide on Airflow assets for data-aware DAG scheduling beyond cron, with cross-team dependency examples.

The guide explains Airflow assets: explicit, visible relationships between DAGs that access the same data, and scheduling on asset updates instead of only time-based methods. An asset may be a table, object-storage file, fine-tuned LLM, or an abstract completed process; the guide covers concepts, basic schedules, updates, UI dependencies, and advanced options.

Based on: Basic asset-based scheduling in Apache Airflow® - Astronomer · Astronomer

HighlightApache Airflow

Airflow 2.4 schedules DAGs on dataset updates, not only time

Apache Airflow 2.4.0 release notes introducing data-aware scheduling via Dataset URIs so producer tasks can trigger consumer DAGs.

Apache Airflow 2.4.0, released 19 September 2022, adds data-aware scheduling (AIP-48): DAGs can schedule on Dataset updates produced by other tasks. Datasets are URI-identified abstracts without direct read/write in this release; the post positions them as a foundation for smaller, chained DAGs and a possible replacement for ExternalTaskSensor or TriggerDagRunOperator in many cases.

Based on: Apache Airflow 2.4.0: That Data Aware Release · Apache Airflow

HighlightAirbyte

Airbyte schema-change policies decide how source drift reaches the destination

Airbyte docs on per-connection schema change detection and propagation: new/removed fields and streams, type changes, and breaking cursor/key cases.

Each Airbyte connection can specify how source schema changes are handled. Cloud checks before sync at most every 15 minutes per source; self-managed at most every 24 hours, with manual refresh available. Behaviors cover new and removed columns and streams, type changes, and immediate pause when a cursor is removed.

Based on: Schema change management | Airbyte Docs · Airbyte