HighlightAmazon Web Services

AWS Glue Data Catalog indexes location, schema, and metrics

AWS documentation on Glue’s central metadata catalog, crawlers, manual tables, and links to Athena, Lake Formation, EMR, and SageMaker.

The AWS Glue Data Catalog is described as a centralized metadata repository that indexes location, schema, and runtime metrics of data sources in metadata tables. Catalog entries can be filled by crawlers that scan internal and external sources, or defined manually. The catalog feeds ETL and integrates with Athena, Lake Formation, EMR, and SageMaker AI.

Based on: Data discovery and cataloging in AWS Glue - AWS Glue · Amazon Web Services

HighlightMicrosoft Learn

Synapse dedicated pool: pick smallest types to shorten rows

Microsoft Learn recommendations for Synapse SQL Dedicated Pool table data types, row length, PolyBase limits, and finding unsupported types when migrating.

The article gives recommendations for defining table data types in Synapse SQL Dedicated Pool, which supports commonly used types listed in CREATE TABLE. Guidance centers on minimizing row length for performance: smallest workable types, tight VARCHAR lengths, VARCHAR over NVARCHAR when possible, capped lengths instead of MAX, and integer types instead of zero-scale decimals. PolyBase external loads cannot exceed 1 MB per row.

Based on: Table data types in Synapse SQL - Azure Synapse Analytics · Microsoft Learn

HighlightMicrosoft Learn

Fabric lakehouse schemas group tables for access and four-part SQL

Microsoft Fabric docs on lakehouse schemas for domain grouping, schema-level access, cross-workspace queries, and schema shortcuts to external Delta.

Lakehouse schemas in Microsoft Fabric group tables into named collections such as sales or hr. Schema-enabled lakehouses support domain browsing, schema-level access with row- and column-level security, four-part workspace.lakehouse.schema.table queries, schema shortcuts to other lakehouses or ADLS Gen2, and features such as materialized lake views. Schema names allow only letters, numbers, and underscores; dbo is the default under Tables.

Based on: Lakehouse schemas - Microsoft Fabric · Microsoft Learn

HighlightDatabricks

Delta table schema evolution: add, reorder, rename, widen types

Databricks guide to explicit and implicit table schema changes, including DDL column operations and conflicts with concurrent writes and streams.

Databricks tables support schema evolution: adding columns at arbitrary positions, reordering, renaming, and type widening. Changes can be made with DDL such as ALTER TABLE or implicitly via DML. Schema updates conflict with concurrent writes, and updating a schema terminates streams reading the table until they are restarted.

Based on: Update table schemas with schema evolution | Databricks on AWS · Databricks

HighlightDatabricks

Databricks Delta constraints: enforced checks vs informational keys

Databricks documentation of NOT NULL and CHECK enforced constraints versus informational primary, foreign, and unique keys on Delta Lake tables.

Databricks supports enforced constraints that reject violating writes and informational primary key, foreign key, and unique constraints that declare relationships without enforcement. All require Delta Lake. Enforced types are NOT NULL and CHECK; adding constraints may raise the table writer protocol and affect external Delta clients.

Based on: Constraints on Databricks | Databricks on AWS · Databricks

HighlightDatabricks

Unity Catalog tags for search, ABAC inheritance, and governed keys

Databricks docs on applying key-value tags to Unity Catalog securable objects, including ABAC inheritance behavior and governed tags.

Tags on Unity Catalog securables are key attributes with optional values used to organize objects and improve workspace search for tables and views. Tag text is stored as plain text and may replicate globally, so sensitive content must not go in tags. For ABAC evaluation, tags on higher objects implicitly apply beneath them but not to columns; governed tags add account-level allowed keys and rules.

Based on: Apply tags to Unity Catalog securable objects | Databricks on AWS · Databricks

HighlightGoogle Cloud

Design BigQuery policy-tag trees around few data classes

Google Cloud best practices for BigQuery policy-tag hierarchies used in column-level access control and dynamic data masking.

Policy tags define access for column-level control and dynamic masking and are presented as an alternative to Resource Manager data governance tags. The guidance is to model a small set of business data classes, map many columns to few tags, and align tags with groups that need different access. Tags can form a tree, often under a root, with an example High/Medium/Low taxonomy and leaf tags such as Credit card and Government ID.

Based on: Best practices for using policy tags in BigQuery | Google Cloud Documentation · Google Cloud

HighlightGoogle Cloud

BigQuery column security via policy and governance tags

Google Cloud intro to BigQuery column-level access control using policy tags or data governance tags, IAM checks at query time, and optional masking.

BigQuery restricts sensitive columns with policy tags from Data Catalog or data governance tags from Resource Manager. Policies are evaluated at query time; optional dynamic masking can replace values with null, default, or hashed content. The workflow is taxonomy and tags, schema annotations on columns, enforce access on the taxonomy, then IAM on each tag, in addition to dataset ACLs.

Based on: Introduction to column-level access control | BigQuery | Google Cloud Documentation · Google Cloud

HighlightSnowflake

Snowflake tags as schema-level key-value labels for governance

Snowflake docs on object tags: schema-level key-value pairs assignable across object types, with inheritance, propagation, and queryable governance use.

A Snowflake tag is a schema-level object stored as a string key-value pair and assigned to other objects. Objects may carry multiple tags; one tag may apply to different object types; values may be shared or unique per assignment. Tags support inheritance down the securable hierarchy, optional automatic propagation, replication of assignments, and centralized or decentralized administration for auditing and reporting.

Based on: Introduction to object tagging | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake classifies columns into semantic and privacy categories

Snowflake Enterprise docs on sensitive data classification, native and custom categories, tags, and Trust Center setup.

Sensitive data classification (Enterprise Edition or higher) automatically discovers sensitive columns and supports governance controls such as tags and masking policies. Each identified column gets a semantic category (e.g., name, national identifier, or custom) and a privacy category (IDENTIFIER, QUASI_IDENTIFIER, or SENSITIVE). Setup and results go through Trust Center; Snowsight can recommend databases likely to hold sensitive data.

Based on: Introduction to sensitive data classification | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake Information Schema is the per-database data dictionary

Snowflake docs on INFORMATION_SCHEMA views and table functions for database and account-level object metadata.

Snowflake’s Information Schema is a SQL-92 ANSI–based data dictionary implemented as a built-in, read-only INFORMATION_SCHEMA in every database. It exposes views for database objects and account-level objects (roles, warehouses, databases) plus table functions for historical and usage data, mixing ANSI-standard and Snowflake-specific views.

Based on: Snowflake Information Schema | Snowflake Documentation · Snowflake

HighlightNeo4j

Neo4j graph modeling ties domain questions to storage shape

Neo4j getting-started overview of graph data modeling steps from domain use cases through test, Cypher load, and refactor.

Graph data modeling defines query logic and stored structure; a well-designed model improves performance, flexibility, and storage use. The process covers understanding the domain and use-case questions, extracting entities and relationships, testing against an initial model, loading test data with Cypher, measuring performance, and refactoring as use cases change.

Based on: What is graph data modeling? - Getting Started · Neo4j