Highlightdbt Labs

Centralize metrics in dbt so every BI tool reads the same definitions

Overview of the dbt Semantic Layer: MetricFlow-backed metrics defined once in the modeling layer for consistent downstream use.

The dbt Semantic Layer, powered by MetricFlow, lets teams define metrics on existing models, handle joins, and expose those definitions to downstream tools. Moving metrics out of BI into the modeling layer keeps business units on the same definitions; a change in dbt refreshes everywhere the metric is invoked. Access permissions control who can use it; Starter or Enterprise-tier accounts are required.

Based on: dbt Semantic Layer | dbt Developer Hub · dbt Labs

Highlightdbt Labs

dbt model contracts: build-time guarantees on column shape before consumers break

dbt Labs docs on model contracts—upfront shape guarantees verified at build, plus governance caveats and support limits.

A dbt model contract is a set of upfront guarantees on the shape of a model’s dataset. At build time dbt checks that the transformation matches the contract or fails. Governance features add trust but can harden rollbacks if adopted too early, and contracts apply only to supported SQL materializations—not Python models, ephemeral models, or several other resource types.

Based on: Model contracts | dbt Developer Hub · dbt Labs

HighlightSalesforce

Salesforce describe API: pull object fields, URLs, and relationships as metadata

REST reference for sObject Describe—full object metadata via GET, with conditional If-Modified-Since headers.

The sObject Describe resource returns complete metadata for a named Salesforce object, including fields, URLs, and child relationships. It is a GET endpoint returning JSON or XML, authenticated with a Bearer token. Optional If-Modified-Since and If-Unmodified-Since headers support conditional responses, including 304 when nothing changed.

Based on: sObject Describe | Reference | REST API Developer Guide | Salesforce Developers · Salesforce

Highlightdbt Labs

MetricFlow: YAML metrics and a semantic graph that generate the SQL

dbt Labs guide introducing MetricFlow as the engine behind the Semantic Layer for defining and querying metrics.

MetricFlow powers dbt’s Semantic Layer with opinionated abstractions for defining and managing metric logic. It builds SQL from YAML (and, from dbt v1.12, Ossie documents), linking semantic models and metrics in a semantic graph. It works with listed warehouses and dbt 1.6+, under Apache 2.0.

Based on: About MetricFlow | dbt Developer Hub · dbt Labs

Highlightmartin.kleppmann.com

Why Thrift, Protobuf, and Avro must survive schema change, not only serialize bytes

Martin Kleppmann on how teams reach schema-based binary formats and why schema evolution is the overlooked requirement.

Kleppmann traces a common path from language-native serialization to JSON to custom binary formats, then to Thrift, Protocol Buffers, or Avro for schema-driven, cross-language encoding. He argues that comparisons often skip what happens when the schema changes. All three formats support evolution so producers and consumers on different versions can still interoperate.

Based on: Schema evolution in Avro, Protocol Buffers and Thrift Martin Kleppmann's blog · martin.kleppmann.com

HighlightConfluent

Compatibility types that keep producers and consumers in sync as schemas change

Confluent Schema Registry docs on schema evolution and backward, forward, full, and transitive compatibility modes.

Schema evolution means changing schemas over time while keeping producers and consumers compatible. Schema Registry compares new versions to prior ones using configurable compatibility types, with BACKWARD as the default. Rules differ by format (Avro, Protobuf, JSON Schema) and by how fields were originally defined.

Based on: Schema Evolution & Compatibility Types | Backward, Forward, Full, Transitive | Confluent Documentation · Confluent

HighlightarXiv

Self-Instruct: Aligning Language Models with Self-Generated Instructions

Introduces Self-Instruct, a framework that improves instruction-following in LLMs by bootstrapping instructions from the model's own generations.

Instruction-tuned language models generalize well zero-shot but depend heavily on limited human-written instruction data. Self-Instruct is a framework that improves instruction-following by bootstrapping off a model's own generations: it generates instructions, inputs, and outputs from the model, filters invalid or similar ones, and uses them to finetune the original model. Applied to vanilla GPT3, it yields a 33% absolute improvement on Super-NaturalInstructions, on par with InstructGPT-001, and leaves only a 5% gap behind it on expert-written novel-task instructions.

Based on: Self-Instruct: Aligning Language Models with Self-Generated Instructions · Annual Meeting of the Association for Computational Linguistics

HighlightarXiv

SciBERT: A Pretrained Language Model for Scientific Text

Releases SciBERT, a BERT-based language model pretrained on scientific text to improve downstream scientific NLP tasks.

Large-scale annotated data for scientific-domain NLP is expensive and hard to obtain. The authors release SciBERT, a pretrained language model based on BERT that leverages unsupervised pretraining on a large multi-domain corpus of scientific publications. SciBERT is evaluated on a suite of tasks including sequence tagging, sentence classification, and dependency parsing, using datasets from a variety of scientific domains. It shows statistically significant improvements over BERT and achieves new state-of-the-art results on several tasks, with code and pretrained models released publicly.

Based on: SciBERT: A Pretrained Language Model for Scientific Text · Conference on Empirical Methods in Natural Language Processing

HighlightarXiv

Efficiently Modeling Long Sequences with Structured State Spaces

Introduces S4, a structured state space sequence model that efficiently captures very long-range dependencies across modalities.

S4 is a sequence model for long-range dependencies where RNNs, CNNs, and Transformers struggle at 10,000+ steps. It builds on the state space model x'=Ax+Bu, y=Cx+Du, and introduces a parameterization that conditions state matrix A with a low-rank correction, enabling stable diagonalization and reducing computation to a Cauchy kernel. This makes prior SSMs far more efficient while preserving their strengths. S4 reaches 91% on sequential CIFAR-10, narrows the gap to Transformers while generating 60x faster, and sets SOTA on Long Range Arena, including the 16k-length Path-X task.

Based on: Efficiently Modeling Long Sequences with Structured State Spaces · International Conference on Learning Representations

HighlightAAAI Conference on Artificial Intelligence

Knowledge Graph Embedding by Translating on Hyperplanes

Proposes TransH, a knowledge graph embedding that models each relation as a hyperplane with a translation to capture complex mapping types.

The paper embeds knowledge graphs of entities and relations into a continuous vector space. The efficient TransE handles reflexive and one-to-many, many-to-one, and many-to-many relations poorly, while richer models fix this but sacrifice efficiency. TransH represents each relation as a hyperplane with a translation on it, preserving these mapping properties at nearly TransE's complexity, plus a sampling trick that cuts false-negative labels. On WordNet and Freebase—link prediction, triplet classification, fact extraction—it significantly improves accuracy over TransE.

Based on: Knowledge Graph Embedding by Translating on Hyperplanes · AAAI Conference on Artificial Intelligence

HighlightarXiv

DeBERTa: Decoding-enhanced BERT with Disentangled Attention

Proposes DeBERTa, improving BERT/RoBERTa with disentangled attention over content and position plus an enhanced mask decoder.

DeBERTa (Decoding-enhanced BERT with disentangled attention) improves BERT and RoBERTa via two techniques. Disentangled attention represents each word with separate content and position vectors, computing attention using disentangled matrices over contents and relative positions. An enhanced mask decoder replaces the output softmax to predict masked tokens in pretraining. Trained on half the data, DeBERTa beats RoBERTa-Large on MNLI (+0.9%), SQuAD v2.0 (+2.3%), and RACE (+3.6%). Code and models are released publicly.

Based on: DeBERTa: Decoding-enhanced BERT with Disentangled Attention · International Conference on Learning Representations

HighlightarXiv

Self-Refine: Iterative Refinement with Self-Feedback

Introduces Self-Refine, a training-free method where one LLM iteratively improves its own outputs via self-generated feedback.

Self-Refine improves LLM outputs through iterative self-feedback and refinement, mirroring how humans revise their writing. A single LLM generates an initial output, then provides feedback on it and uses that feedback to refine itself, requiring no supervised data, extra training, or reinforcement learning. Evaluated across 7 diverse tasks on GPT-3.5, ChatGPT, and GPT-4, Self-Refine outputs are preferred by humans and automatic metrics, improving performance by about 20% absolute on average. The work shows even top LLMs like GPT-4 can be improved at test time.

Based on: Self-Refine: Iterative Refinement with Self-Feedback · Neural Information Processing Systems