HighlightAmazon Web Services

AWS Glue Data Catalog indexes location, schema, and metrics

AWS documentation on Glue’s central metadata catalog, crawlers, manual tables, and links to Athena, Lake Formation, EMR, and SageMaker.

The AWS Glue Data Catalog is described as a centralized metadata repository that indexes location, schema, and runtime metrics of data sources in metadata tables. Catalog entries can be filled by crawlers that scan internal and external sources, or defined manually. The catalog feeds ETL and integrates with Athena, Lake Formation, EMR, and SageMaker AI.

Based on: Data discovery and cataloging in AWS Glue - AWS Glue · Amazon Web Services

HighlightMicrosoft Learn

Fabric lakehouse schemas group tables for access and four-part SQL

Microsoft Fabric docs on lakehouse schemas for domain grouping, schema-level access, cross-workspace queries, and schema shortcuts to external Delta.

Lakehouse schemas in Microsoft Fabric group tables into named collections such as sales or hr. Schema-enabled lakehouses support domain browsing, schema-level access with row- and column-level security, four-part workspace.lakehouse.schema.table queries, schema shortcuts to other lakehouses or ADLS Gen2, and features such as materialized lake views. Schema names allow only letters, numbers, and underscores; dbo is the default under Tables.

Based on: Lakehouse schemas - Microsoft Fabric · Microsoft Learn

HighlightMicrosoft Learn

Purview and Fabric as one path from source to Power BI lineage

Microsoft Learn article on using Microsoft Purview with Microsoft Fabric for estate-wide governance, classification, and end-to-end lineage.

Microsoft Purview and Microsoft Fabric are presented as parts of the Microsoft intelligent data platform for storing, analyzing, and governing data together. Combined, they are said to cover the estate and lineage from data source to Power BI report without stitching multiple vendors. Purview is described as governance, risk, and compliance coverage across Microsoft 365, on-premises, multicloud, and SaaS.

Based on: Use Microsoft Purview to Govern Microsoft Fabric - Microsoft Fabric · Microsoft Learn

HighlightGoogle Cloud

Knowledge Catalog: Gemini context graph for agents and governance

Google Cloud overview of Knowledge Catalog (formerly Dataplex Universal Catalog): an AI-powered catalog that builds a context graph to ground agents.

As of April 10, 2026, Dataplex Universal Catalog is renamed Knowledge Catalog while API, client library, CLI, and IAM names stay the same. The product is described as a Gemini-powered catalog that extracts semantics from structured and unstructured data into a dynamic context graph for discovery, policy, and agent grounding to reduce hallucinations.

Based on: Knowledge Catalog overview | Google Cloud Documentation · Google Cloud

HighlightDatabricks

Databricks Delta constraints: enforced checks vs informational keys

Databricks documentation of NOT NULL and CHECK enforced constraints versus informational primary, foreign, and unique keys on Delta Lake tables.

Databricks supports enforced constraints that reject violating writes and informational primary key, foreign key, and unique constraints that declare relationships without enforcement. All require Delta Lake. Enforced types are NOT NULL and CHECK; adding constraints may raise the table writer protocol and affect external Delta clients.

Based on: Constraints on Databricks | Databricks on AWS · Databricks

HighlightDatabricks

Unity Catalog tags for search, ABAC inheritance, and governed keys

Databricks docs on applying key-value tags to Unity Catalog securable objects, including ABAC inheritance behavior and governed tags.

Tags on Unity Catalog securables are key attributes with optional values used to organize objects and improve workspace search for tables and views. Tag text is stored as plain text and may replicate globally, so sensitive content must not go in tags. For ABAC evaluation, tags on higher objects implicitly apply beneath them but not to columns; governed tags add account-level allowed keys and rules.

Based on: Apply tags to Unity Catalog securable objects | Databricks on AWS · Databricks

HighlightGoogle Cloud

Design BigQuery policy-tag trees around few data classes

Google Cloud best practices for BigQuery policy-tag hierarchies used in column-level access control and dynamic data masking.

Policy tags define access for column-level control and dynamic masking and are presented as an alternative to Resource Manager data governance tags. The guidance is to model a small set of business data classes, map many columns to few tags, and align tags with groups that need different access. Tags can form a tree, often under a root, with an example High/Medium/Low taxonomy and leaf tags such as Credit card and Government ID.

Based on: Best practices for using policy tags in BigQuery | Google Cloud Documentation · Google Cloud

HighlightGoogle Cloud

BigQuery column security via policy and governance tags

Google Cloud intro to BigQuery column-level access control using policy tags or data governance tags, IAM checks at query time, and optional masking.

BigQuery restricts sensitive columns with policy tags from Data Catalog or data governance tags from Resource Manager. Policies are evaluated at query time; optional dynamic masking can replace values with null, default, or hashed content. The workflow is taxonomy and tags, schema annotations on columns, enforce access on the taxonomy, then IAM on each tag, in addition to dataset ACLs.

Based on: Introduction to column-level access control | BigQuery | Google Cloud Documentation · Google Cloud

HighlightREDCap Consortium / Vanderbilt

REDCap: consortium survey and database software with FHIR hooks

REDCap Consortium page describing the secure web app for research surveys and databases, including data dictionary design, exports, compliance, and FHIR.

REDCap is a secure web application for building and managing online surveys and databases, free to REDCap Consortium partners. Design can use an Online Designer or an Excel data dictionary upload; features include audit trails, multi-site access, statistical exports, and installs aimed at HIPAA, 21 CFR Part 11, and FISMA. Structured EHR data can be pulled via FHIR with OAuth2 through Clinical Data Interoperability Services.

Based on: Software – REDCap · REDCap Consortium / Vanderbilt

HighlightSnowflake

Semantic views put business metrics and entities in the database

Snowflake overview of Semantic Views: schema-level objects that define metrics, entities, and relationships for consistent business meaning and Cortex Agents.

A Semantic View is a schema-level Snowflake object that stores business metrics, entities, and relationships as metadata atop physical data. It supplies consistent definitions across applications, is queryable with SELECT, usable by Cortex Agents, and shareable via listings. The docs argue this layer fixes the gap between business language and opaque column names and stops inconsistent metric calculations across reports.

Based on: Overview of semantic views | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake tags as schema-level key-value labels for governance

Snowflake docs on object tags: schema-level key-value pairs assignable across object types, with inheritance, propagation, and queryable governance use.

A Snowflake tag is a schema-level object stored as a string key-value pair and assigned to other objects. Objects may carry multiple tags; one tag may apply to different object types; values may be shared or unique per assignment. Tags support inheritance down the securable hierarchy, optional automatic propagation, replication of assignments, and centralized or decentralized administration for auditing and reporting.

Based on: Introduction to object tagging | Snowflake Documentation · Snowflake

HighlightSnowflake

Snowflake classifies columns into semantic and privacy categories

Snowflake Enterprise docs on sensitive data classification, native and custom categories, tags, and Trust Center setup.

Sensitive data classification (Enterprise Edition or higher) automatically discovers sensitive columns and supports governance controls such as tags and masking policies. Each identified column gets a semantic category (e.g., name, national identifier, or custom) and a privacy category (IDENTIFIER, QUASI_IDENTIFIER, or SENSITIVE). Setup and results go through Trust Center; Snowsight can recommend databases likely to hold sensitive data.

Based on: Introduction to sensitive data classification | Snowflake Documentation · Snowflake