HighlightSnowflake

Snowflake tags as schema-level key-value labels for governance

Snowflake docs on object tags: schema-level key-value pairs assignable across object types, with inheritance, propagation, and queryable governance use.

A Snowflake tag is a schema-level object stored as a string key-value pair and assigned to other objects. Objects may carry multiple tags; one tag may apply to different object types; values may be shared or unique per assignment. Tags support inheritance down the securable hierarchy, optional automatic propagation, replication of assignments, and centralized or decentralized administration for auditing and reporting.

Based on: Introduction to object tagging | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake classifies columns into semantic and privacy categories

Snowflake Enterprise docs on sensitive data classification, native and custom categories, tags, and Trust Center setup.

Sensitive data classification (Enterprise Edition or higher) automatically discovers sensitive columns and supports governance controls such as tags and masking policies. Each identified column gets a semantic category (e.g., name, national identifier, or custom) and a privacy category (IDENTIFIER, QUASI_IDENTIFIER, or SENSITIVE). Setup and results go through Trust Center; Snowsight can recommend databases likely to hold sensitive data.

Based on: Introduction to sensitive data classification | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake Information Schema is the per-database data dictionary

Snowflake docs on INFORMATION_SCHEMA views and table functions for database and account-level object metadata.

Snowflake’s Information Schema is a SQL-92 ANSI–based data dictionary implemented as a built-in, read-only INFORMATION_SCHEMA in every database. It exposes views for database objects and account-level objects (roles, warehouses, databases) plus table functions for historical and usage data, mixing ANSI-standard and Snowflake-specific views.

Based on: Snowflake Information Schema | Snowflake Documentation · Snowflake

HighlightNeo4j

Neo4j graph modeling ties domain questions to storage shape

Neo4j getting-started overview of graph data modeling steps from domain use cases through test, Cypher load, and refactor.

Graph data modeling defines query logic and stored structure; a well-designed model improves performance, flexibility, and storage use. The process covers understanding the domain and use-case questions, extracting entities and relationships, testing against an initial model, loading test data with Cypher, measuring performance, and refactoring as use cases change.

Based on: What is graph data modeling? - Getting Started · Neo4j

HighlightNeo4j

Neo4j constraints enforce uniqueness, existence, type, and keys

Neo4j Cypher manual listing property uniqueness, existence, type, and key constraints, and preferring graph types for schema.

Neo4j provides property uniqueness, existence (Enterprise), type (Enterprise), and key (Enterprise) constraints on nodes by label or relationships by type. Older CREATE CONSTRAINT syntax still adds constraints to a database’s graph type, but the docs recommend defining schema via graph types for richer constraint kinds and simpler long-term maintenance.

Based on: Constraints - Cypher Manual · Neo4j

HighlightNeo4j

Neo4j MERGE: match-or-create patterns with ON MATCH/ON CREATE

Neo4j Cypher manual for MERGE, which binds existing patterns or creates missing ones, with constraint guidance.

MERGE combines MATCH and CREATE: if the exact pattern exists it binds like MATCH; otherwise it creates like CREATE. ON MATCH and ON CREATE allow different actions for each case. The manual recommends creating constraints before merging for index-backed performance and to prevent unintended divergent data.

Based on: MERGE - Cypher Manual · Neo4j

HighlightAstronomer

Astronomer on OpenLineage and Airflow for pipeline lineage

Astronomer Learn guide covering data lineage concepts and why integrating lineage with Apache Airflow matters.

The guide defines data lineage as tracking and visualizing data from origin through downstream consumption, and lists uses such as understanding sources, troubleshooting failures, managing PII, and regulatory compliance. It positions Airflow as a central orchestrator for integrating lineage and introduces how lineage works with Airflow, with pointers to Astro Observe and related webinars.

Based on: Integrate OpenLineage and Airflow - Astronomer · Astronomer

HighlightAstronomer

Astronomer guide: schedule Airflow DAGs on asset updates

Astronomer Learn guide on Airflow assets for data-aware DAG scheduling beyond cron, with cross-team dependency examples.

The guide explains Airflow assets: explicit, visible relationships between DAGs that access the same data, and scheduling on asset updates instead of only time-based methods. An asset may be a table, object-storage file, fine-tuned LLM, or an abstract completed process; the guide covers concepts, basic schedules, updates, UI dependencies, and advanced options.

Based on: Basic asset-based scheduling in Apache Airflow® - Astronomer · Astronomer

HighlightApache Airflow

Airflow 2.4 schedules DAGs on dataset updates, not only time

Apache Airflow 2.4.0 release notes introducing data-aware scheduling via Dataset URIs so producer tasks can trigger consumer DAGs.

Apache Airflow 2.4.0, released 19 September 2022, adds data-aware scheduling (AIP-48): DAGs can schedule on Dataset updates produced by other tasks. Datasets are URI-identified abstracts without direct read/write in this release; the post positions them as a foundation for smaller, chained DAGs and a possible replacement for ExternalTaskSensor or TriggerDagRunOperator in many cases.

Based on: Apache Airflow 2.4.0: That Data Aware Release · Apache Airflow

HighlightGoogle Search Central

Google Search uses on-page structured data as explicit meaning for rich results

Google Search Central intro: structured data markup gives Search explicit page meaning, enables rich results, and cites publisher case studies.

Google Search can use structured data on a page as explicit clues about content—for example recipe ingredients, time, and calories. That markup can enable rich results. Google also uses it to understand page content and gather facts about entities such as people, books, and companies. Case studies cite CTR and engagement lifts for several publishers.

Based on: Intro to How Structured Data Markup Works | Google Search Central | Documentation | Google for Developers · Google Search Central

HighlightAirbyte

Airbyte schema-change policies decide how source drift reaches the destination

Airbyte docs on per-connection schema change detection and propagation: new/removed fields and streams, type changes, and breaking cursor/key cases.

Each Airbyte connection can specify how source schema changes are handled. Cloud checks before sync at most every 15 minutes per source; self-managed at most every 24 hours, with manual refresh available. Behaviors cover new and removed columns and streams, type changes, and immediate pause when a cursor is removed.

Based on: Schema change management | Airbyte Docs · Airbyte

HighlightAirbyte

Airbyte Protocol standardizes source and destination actors for ELT pipelines

Airbyte docs defining the protocol: actors, catalog/stream/field primitives, and STDIO JSON vs Socket Protobuf data channels.

The Airbyte Protocol specifies standard components and interactions that declare an ELT pipeline. Sources and destinations are actors with standard interfaces; data is described via catalog, configured catalog, stream, configured stream, and field. Two channel modes exist: STDIO with serialized JSON messages, and Socket mode with Protocol Buffers over Unix domain sockets.

Based on: Airbyte Protocol | Airbyte Docs · Airbyte