HighlightMicrosoft Learn

Synapse dedicated pool: pick smallest types to shorten rows

Microsoft Learn recommendations for Synapse SQL Dedicated Pool table data types, row length, PolyBase limits, and finding unsupported types when migrating.

The article gives recommendations for defining table data types in Synapse SQL Dedicated Pool, which supports commonly used types listed in CREATE TABLE. Guidance centers on minimizing row length for performance: smallest workable types, tight VARCHAR lengths, VARCHAR over NVARCHAR when possible, capped lengths instead of MAX, and integer types instead of zero-scale decimals. PolyBase external loads cannot exceed 1 MB per row.

Based on: Table data types in Synapse SQL - Azure Synapse Analytics · Microsoft Learn

HighlightDatabricks

Delta table schema evolution: add, reorder, rename, widen types

Databricks guide to explicit and implicit table schema changes, including DDL column operations and conflicts with concurrent writes and streams.

Databricks tables support schema evolution: adding columns at arbitrary positions, reordering, renaming, and type widening. Changes can be made with DDL such as ALTER TABLE or implicitly via DML. Schema updates conflict with concurrent writes, and updating a schema terminates streams reading the table until they are restarted.

Based on: Update table schemas with schema evolution | Databricks on AWS · Databricks

HighlightDatabricks

Databricks Delta constraints: enforced checks vs informational keys

Databricks documentation of NOT NULL and CHECK enforced constraints versus informational primary, foreign, and unique keys on Delta Lake tables.

Databricks supports enforced constraints that reject violating writes and informational primary key, foreign key, and unique constraints that declare relationships without enforcement. All require Delta Lake. Enforced types are NOT NULL and CHECK; adding constraints may raise the table writer protocol and affect external Delta clients.

Based on: Constraints on Databricks | Databricks on AWS · Databricks

HighlightApache Airflow

Airflow 2.4 schedules DAGs on dataset updates, not only time

Apache Airflow 2.4.0 release notes introducing data-aware scheduling via Dataset URIs so producer tasks can trigger consumer DAGs.

Apache Airflow 2.4.0, released 19 September 2022, adds data-aware scheduling (AIP-48): DAGs can schedule on Dataset updates produced by other tasks. Datasets are URI-identified abstracts without direct read/write in this release; the post positions them as a foundation for smaller, chained DAGs and a possible replacement for ExternalTaskSensor or TriggerDagRunOperator in many cases.

Based on: Apache Airflow 2.4.0: That Data Aware Release · Apache Airflow

HighlightAirbyte

Airbyte schema-change policies decide how source drift reaches the destination

Airbyte docs on per-connection schema change detection and propagation: new/removed fields and streams, type changes, and breaking cursor/key cases.

Each Airbyte connection can specify how source schema changes are handled. Cloud checks before sync at most every 15 minutes per source; self-managed at most every 24 hours, with manual refresh available. Behaviors cover new and removed columns and streams, type changes, and immediate pause when a cursor is removed.

Based on: Schema change management | Airbyte Docs · Airbyte

HighlightAirbyte

Airbyte Protocol standardizes source and destination actors for ELT pipelines

Airbyte docs defining the protocol: actors, catalog/stream/field primitives, and STDIO JSON vs Socket Protobuf data channels.

The Airbyte Protocol specifies standard components and interactions that declare an ELT pipeline. Sources and destinations are actors with standard interfaces; data is described via catalog, configured catalog, stream, configured stream, and field. Two channel modes exist: STDIO with serialized JSON messages, and Socket mode with Protocol Buffers over Unix domain sockets.

Based on: Airbyte Protocol | Airbyte Docs · Airbyte

HighlightW3C Shape Expressions Community Group / shex.io

ShEx 2.1 defines shapes and node constraints for describing RDF graph structure

Final Community Group Report (8 Oct 2019) for Shape Expressions Language 2.1: RDF node and graph structure descriptions for validation and interfaces.

Shape Expressions (ShEx) describe RDF nodes and graph structures: node constraints on IRIs, blank nodes, or literals, and shapes over triples with predicates, cardinalities, and datatypes. Shapes can communicate structures for processes or interfaces, generate or validate data, or drive user interfaces. ShEx 2.1 adds IMPORTS and language tag value sets.

Based on: Shape Expressions Language 2.1 · W3C Shape Expressions Community Group / shex.io

HighlightOracle MySQL

MySQL CHECK constraints encode boolean row rules with optional enforcement

MySQL 8.4 reference for CREATE TABLE CHECK constraints: naming, boolean expr, ENFORCED/NOT ENFORCED, table vs column form.

MySQL 8.4 permits core table and column CHECK constraints for all storage engines. A CHECK (expr) must evaluate to TRUE or UNKNOWN per row; FALSE is a violation. Constraints may be ENFORCED or NOT ENFORCED, named with an optional symbol, and declared as table- or column-level.

Based on: MySQL :: MySQL 8.4 Reference Manual :: 15.1.20.6 CHECK Constraints · Oracle MySQL

HighlightApache Software Foundation

Iceberg evolves schema and partitions in place without rewriting data

Apache Iceberg docs on metadata-only schema changes and partition evolution that leave existing files intact.

Iceberg supports in-place table evolution for schemas—including nested structures—and for partition layouts as data volume changes, without rewriting data or migrating tables. Schema operations include add, drop, rename, widen types, and reorder; updates are metadata changes tracked by unique column IDs with correctness guarantees. Partition specs can change while old data keeps its prior layout.

Based on: Evolution - Apache Iceberg™ · Apache Software Foundation

HighlightJSON Schema (IETF-oriented draft)

JSON Schema Validation vocabulary asserts structure, meaning, and UI hints

Internet-Draft specifying the JSON Schema vocabulary for instance validation, document meaning, and UI hints.

This draft specifies a vocabulary for JSON Schema used in JSON instance validation. It describes meanings of JSON documents, hints for user interfaces, and assertions about valid document shape. Authors are Wright, Andrews, and Hutton; the draft was published 16 June 2022 as informational IETF work in progress.

Based on: JSON Schema Validation: A Vocabulary for Structural Validation of JSON · JSON Schema (IETF-oriented draft)

HighlightJSON Schema (IETF-oriented draft)

JSON Schema defines a media type for describing JSON document structure

IETF-oriented Internet-Draft for JSON Schema as application/schema+json, including instance media type notes.

This Internet-Draft defines JSON Schema as the media type application/schema+json, a JSON-based format for describing JSON structure, extraction, and interaction. It also discusses application/schema-instance+json for richer integration than plain application/json. The draft is informational, authored by Wright, Andrews, Hutton, and Dennis, published 16 June 2022.

Based on: JSON Schema: A Media Type for Describing JSON Documents · JSON Schema (IETF-oriented draft)

HighlightGoogle

Proto3 defines typed messages and stable field numbers for shared payloads

Google’s proto3 language guide on .proto syntax, message fields, types, and generating data access classes.

The proto3 Language Guide explains how to structure protocol buffer data with .proto syntax and how generated access classes are produced. A message declares named, typed fields each assigned a field number. The edition or syntax line must be first; if omitted, the compiler assumes proto2.

Based on: Language Guide (proto 3) · Google