HighlightAnthropic

Anthropic open-sources MCP as a universal AI-to-data connection standard

Anthropic announces the Model Context Protocol for two-way links between AI tools and data sources via servers and clients.

On 25 Nov 2024 Anthropic open-sourced MCP to connect AI assistants to content repos, business tools, and development environments under one protocol instead of per-source integrations. Release includes the spec and SDKs, local server support in Claude Desktop, an open-source server repo, and pre-built servers for systems such as Google Drive, Slack, GitHub, Git, Postgres, and Puppeteer.

Based on: Introducing the Model Context Protocol · Anthropic

Highlightdbt Labs

dbt measures are column aggregations, now migrating to simple metrics

dbt documentation for measures as aggregations on model columns, with deprecation toward type: simple metrics.

Measures aggregate columns and can stand alone or feed complex metrics. Parameters include name, agg (sum, max, min, average, median, count_distinct, percentile, sum_boolean), expr, non_additive_dimension, and related fields. The new spec deprecates measures in favor of simple metrics under metrics.

Based on: Measures | dbt Developer Hub · dbt Labs

Highlightdbt Labs

MetricFlow centralizes metric specs and SQL construction in dbt

dbt intro to defining metrics with MetricFlow as the Semantic Layer component that builds SQL and enforces consistency.

MetricFlow in dbt centrally defines metrics and builds SQL from semantic models and metric specs. It aims to cut duplicative coding, support governance of company metrics, and keep consumer results consistent; defined metrics can be queried in development and, on higher plans, in downstream tools.

Based on: Build your metrics | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt semantic models are MetricFlow graph nodes tied to DAG models

dbt docs explain semantic models as YAML-configured nodes and entities that MetricFlow uses for the Semantic Layer.

In dbt v1.12+, semantic models are the foundation for MetricFlow’s Semantic Layer: nodes linked by entities in a semantic graph, annotated on dbt models for metric use. Each DAG model maps to one semantic_model YAML block (or Apache Ossie docs); components include name, time dimension, and entities.

Based on: Semantic models | dbt Developer Hub · dbt Labs

HighlightConfluent

Confluent Schema Registry centralizes Kafka schema validate-and-evolve rules

Confluent documentation overview of Schema Registry for managing, validating, and evolving schemas on Kafka topics.

Schema Registry is a centralized store for topic message schemas and for serialize/deserialize over the network. Producers and consumers use it for consistency and compatibility as schemas change; Confluent frames it as a data-governance component covering quality, standards, lineage visibility, audit, and collaboration, on Cloud and Platform.

Based on: Schema Registry for Confluent Platform | Confluent Documentation · Confluent

HighlightBitol (LF AI & Data)

ODCS v3.2.0 standardizes the sections of a producer–consumer data contract

Bitol’s Open Data Contract Standard defines a YAML-oriented structure for agreements between data producers and consumers.

The Open Data Contract Standard (ODCS) v3.2.0, under Apache 2.0 from Bitol at LF AI & Data, describes how to structure a data contract. Contracts cover schema, quality, SLA, roles, infrastructure, and related sections, with JSON Schema for YAML validation and media type application/odcs+yaml;version=3.2.0.

Based on: GitHub - bitol-io/open-data-contract-standard: Home of the Open Data Contract Standard (ODCS). · Bitol (LF AI & Data)

HighlightGoogle Cloud

LookML puts join structure and query content in one reusable semantic model

Google Cloud docs introduce LookML as Looker’s modeling language for dimensions, aggregates, joins, and SQL generation.

LookML (Looker Modeling Language) is a dependency-style language for building semantic data models over SQL databases. Models and views define joins and calculations once; Looker generates dialect-independent SQL so analysts avoid repeating expressions and business users query without writing SQL.

Based on: Introduction to LookML | Looker | Google Cloud Documentation · Google Cloud

HighlightDatabricks

Unity Catalog: one governance layer under Databricks data and AI assets

Databricks docs on Unity Catalog—access control, lineage, and auditing for data and AI via a three-level object model.

Unity Catalog is Databricks’ unified governance layer for data and AI. When enabled, it enforces access control, tracks lineage, and logs activity under workspace interactions. Assets are securable objects in a catalog.schema.object namespace; tables and volumes may be managed or external. An open-source implementation also exists.

Based on: What is Unity Catalog? | Databricks on AWS · Databricks

HighlightSnowflake

Cortex Analyst: managed text-to-SQL over Snowflake structure via REST

Snowflake docs on Cortex Analyst—an LLM feature for natural-language questions on structured data, exposed as a REST API.

Cortex Analyst is a fully managed Snowflake Cortex feature that answers business questions on structured Snowflake data in natural language without users writing SQL. It ships as a REST API for embedding in apps and aims to produce accurate text-to-SQL without teams building custom RAG stacks or managing GPUs. Snowflake recommends transitioning to Cortex Agents, which includes Analyst capabilities.

Based on: Cortex Analyst | Snowflake Documentation · Snowflake

Highlightmartinfowler.com

Data mesh: domain-owned data products instead of a central lake monolith

Zhamak Dehghani’s 2019 essay on moving from monolithic data lakes to a distributed data mesh.

Dehghani argues enterprise data lakes often fail at scale through centralization, coupled pipelines, and siloed ownership. The proposed shift treats domains as first-class, applies platform thinking for self-serve infrastructure, and treats data as a product with discoverability, addressability, trust, self-describing semantics, interoperability, and secure access.

Based on: How to Move Beyond a Monolithic Data Lake to a Distributed Data Mesh · martinfowler.com

HighlightCube Dev

Cube’s semantic layer as the shared surface for BI, embed, and AI agents

Cube docs: agentic analytics on an open-source semantic layer for internal BI and embedded analytics.

Cube is positioned as an agentic analytics platform on a semantic layer for internal BI and embedded analytics. Cube Core centralizes metrics, joins, access rules, and caching upstream of BI tools, apps, and agents. Agents query via Semantic SQL through the semantic layer runtime with validation and access policies, not by writing free-form warehouse SQL.

Based on: Introduction - Cube Documentation · Cube Dev

HighlightPact Foundation

Pact contract tests: verify service messages without burning the house down

Pact introduction: code-first contract testing for HTTP and message integrations between services.

Pact is a code-first tool for testing HTTP and message integrations with contract tests. Contract tests check that messages between applications match a shared understanding documented in a contract, as an alternative to expensive, brittle end-to-end integration tests. The approach fits especially well when many services must communicate.

Based on: Introduction | Pact Docs · Pact Foundation