HighlightarXiv

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Introduces the Switch Transformer, a simplified sparse mixture-of-experts model that scales to trillion parameters at constant compute cost.

Unlike standard models that reuse the same parameters for all inputs, Mixture of Experts (MoE) selects different parameters per example, giving huge, sparsely activated models at constant compute. Adoption has been limited by complexity, communication cost, and training instability, which the Switch Transformer addresses by simplifying routing and reducing overheads. New techniques tame instabilities and enable bfloat16 training. Based on T5, it delivers up to 7x faster pre-training, gains across 101 languages, and trillion-parameter models with a 4x speedup over T5-XXL.

Based on: Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity · Journal of machine learning research

HighlightarXiv

Sparks of Artificial General Intelligence: Early experiments with GPT-4

Investigates an early version of GPT-4, arguing it shows more general intelligence than prior models across many domains and tasks.

The paper reports an investigation of an early, still-in-development version of OpenAI's GPT-4, trained at unprecedented scale. The authors argue it belongs to a new cohort of LLMs with more general intelligence than earlier AI, solving novel, hard tasks across mathematics, coding, vision, medicine, law, and psychology without special prompting, often near or beyond human level. They suggest it is an early, incomplete form of AGI, stress its limitations, and discuss challenges ahead, including moving beyond next-word prediction.

Based on: Sparks of Artificial General Intelligence: Early experiments with GPT-4 · arXiv.org

HighlightarXiv

Pointer Sentinel Mixture Models

Introduces the pointer sentinel mixture architecture that lets neural sequence models copy words from recent context or use a softmax classifier.

Softmax neural sequence models reach top language modeling only with large hidden states and vocabularies, yet still fail on rare or unseen words even when context is unambiguous. The authors propose a pointer sentinel mixture that either reproduces a word from recent context or generates one via a standard softmax. Their pointer sentinel-LSTM reaches 70.9 perplexity on Penn Treebank using far fewer parameters, and they release the WikiText corpus for evaluating longer contexts and larger vocabularies.

Based on: Pointer Sentinel Mixture Models · International Conference on Learning Representations

HighlightarXiv

Generating Sequences With Recurrent Neural Networks

Shows how LSTM recurrent networks generate complex sequences with long-range structure by predicting one data point at a time.

Long Short-term Memory recurrent neural networks can generate complex sequences with long-range structure simply by predicting one data point at a time. The approach is demonstrated on discrete text data and real-valued online handwriting. It is then extended to handwriting synthesis by letting the network condition its predictions on a text sequence, producing highly realistic cursive handwriting across a wide variety of styles.

Based on: Generating Sequences With Recurrent Neural Networks · arXiv.org

HighlightarXiv

SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing

Describes SentencePiece, a language-independent subword tokenizer and detokenizer that trains directly from raw sentences for neural text processing.

SentencePiece is a language-independent subword tokenizer and detokenizer built for neural text processing, including neural machine translation. Unlike existing tools that assume pre-tokenized word sequences as input, it can train subword models directly from raw sentences, enabling a fully end-to-end, language-independent pipeline. In an English-Japanese NMT experiment, training directly from raw sentences matches the accuracy of direct subword training. Open-source C++ and Python implementations are released under the Apache 2 license.

Based on: SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing · Conference on Empirical Methods in Natural Language Processing

HighlightarXiv

OPT: Open Pre-trained Transformer Language Models

Releases OPT, a suite of open decoder-only pretrained transformers from 125M to 175B parameters, matching GPT-3.

Large language models trained for hundreds of thousands of compute days show strong zero- and few-shot ability, but their cost makes them hard to replicate, and available ones expose no full weights for study. OPT (Open Pre-trained Transformers) is a suite of decoder-only pretrained transformers from 125M to 175B parameters, shared fully and responsibly with researchers. OPT-175B is comparable to GPT-3 while requiring only 1/7th the carbon footprint to develop. The authors also release a logbook of infrastructure challenges and code for experimenting with the models.

Based on: OPT: Open Pre-trained Transformer Language Models · arXiv.org

HighlightarXiv

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

FlashAttention: an IO-aware exact attention algorithm using tiling to cut GPU memory traffic, speeding Transformer training and enabling longer context.

Self-attention scales quadratically with sequence length, making Transformers slow on long sequences, and approximate methods cut compute but often lose quality without speedups. The authors argue the missing principle is IO-awareness, counting reads and writes between GPU memory levels. FlashAttention is an IO-aware exact attention algorithm that uses tiling to cut transfers between GPU HBM and on-chip SRAM, needing fewer HBM accesses. It trains Transformers faster (15% on BERT-large, 3x on GPT-2, 2.4x on long-range arena) and enables longer context and new capabilities.

Based on: FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness · Neural Information Processing Systems

HighlightarXiv

Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings

Shows word embeddings encode gender stereotypes and proposes a geometric method to remove them while preserving useful structure.

Word embeddings, even those trained on Google News, encode female/male gender stereotypes to a disturbing degree, and their widespread use tends to amplify these biases. Geometrically, gender bias is captured by a direction in the space, and gender-neutral words are linearly separable from gender-definitional words. The authors use these properties to modify embeddings, removing stereotypical associations like receptionist-female while keeping legitimate ones like queen-female. Evaluations show significant bias reduction while preserving clustering and analogy-solving ability.

Based on: Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings · Neural Information Processing Systems

HighlightarXiv

Bidirectional LSTM-CRF Models for Sequence Tagging

Applies LSTM, BI-LSTM, LSTM-CRF, and BI-LSTM-CRF models to sequence tagging, reaching state-of-the-art on POS, chunking, and NER.

This paper proposes a range of LSTM-based models for sequence tagging, including LSTM, bidirectional LSTM, LSTM with a CRF layer, and bidirectional LSTM with a CRF layer (BI-LSTM-CRF). It is the first work to apply a BI-LSTM-CRF model to NLP sequence tagging benchmarks. The model uses both past and future input features via the bidirectional LSTM and sentence-level tag information via the CRF layer, producing state-of-the-art or comparable accuracy on POS tagging, chunking, and NER while being robust and less reliant on word embeddings.

Based on: Bidirectional LSTM-CRF Models for Sequence Tagging · arXiv.org

HighlightarXiv

Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Introduces Tree of Thoughts, a framework letting LLMs explore and self-evaluate multiple reasoning paths with lookahead and backtracking.

LLMs are widely used for problem solving but remain confined to token-level, left-to-right inference, limiting tasks needing exploration or strategic lookahead. The authors introduce Tree of Thoughts (ToT), generalizing Chain-of-Thought prompting by letting models explore coherent text units ('thoughts') as intermediate steps. ToT considers multiple reasoning paths, self-evaluates choices, and looks ahead or backtracks for global decisions. On Game of 24, Creative Writing, and Mini Crosswords it sharply improves results—raising GPT-4's Game of 24 success from 4% to 74%.

Based on: Tree of Thoughts: Deliberate Problem Solving with Large Language Models · Neural Information Processing Systems

HighlightarXiv

Transformer-XL: Attentive Language Models beyond a Fixed-Length Context

Proposes Transformer-XL, adding segment-level recurrence and a new positional encoding so language models learn dependency beyond a fixed-length context.

Transformers learn long-term dependency but are limited by a fixed-length context in language modeling. The authors propose Transformer-XL, combining segment-level recurrence with a novel positional encoding to learn dependencies beyond a fixed length without disrupting temporal coherence, also resolving context fragmentation. It captures dependency 80% longer than RNNs and 450% longer than vanilla Transformers, performs better on short and long sequences, is up to 1,800x faster at evaluation, and sets new state-of-the-art results on five benchmarks.

Based on: Transformer-XL: Attentive Language Models beyond a Fixed-Length Context · Annual Meeting of the Association for Computational Linguistics

HighlightarXiv

HuggingFace's Transformers: State-of-the-art Natural Language Processing

Presents Transformers, an open-source library offering unified access to state-of-the-art Transformer architectures and curated pretrained models.

Recent NLP progress has been driven by advances in Transformer architectures, which enable higher-capacity models, and by pretraining, which lets that capacity be used across many tasks. Transformers is an open-source library that brings these advances to the broader machine learning community, providing state-of-the-art Transformer architectures under a unified API together with a curated collection of community-contributed pretrained models. It is designed to be extensible for researchers, simple for practitioners, and fast and robust in industrial deployment.

Based on: HuggingFace's Transformers: State-of-the-art Natural Language Processing · arXiv.org