HighlightarXiv

Modeling Relational Data with Graph Convolutional Networks

Introduces R-GCNs, graph convolutional networks for multi-relational knowledge bases, applied to link prediction and entity classification.

Knowledge graphs remain incomplete even at their largest (Yago, DBpedia, Wikidata). The authors introduce Relational Graph Convolutional Networks (R-GCNs) for two knowledge base completion tasks: link prediction, recovering missing subject-predicate-object triples, and entity classification, recovering missing attributes. R-GCNs extend graph neural networks to highly multi-relational data. Effective stand-alone for entity classification, an R-GCN encoder also improves factorization models like DistMult, giving a 29.8% gain on FB15k-237 over a decoder-only baseline.

Based on: Modeling Relational Data with Graph Convolutional Networks · Extended Semantic Web Conference

HighlightarXiv

The Power of Scale for Parameter-Efficient Prompt Tuning

Introduces prompt tuning, learning soft prompts via backpropagation to adapt frozen language models, matching full model tuning as scale grows.

The paper explores prompt tuning, a mechanism for learning soft prompts that condition frozen language models for downstream tasks. Unlike GPT-3's discrete text prompts, soft prompts are learned via backpropagation from labeled examples and outperform GPT-3's few-shot learning by a large margin. Ablations with T5 show the method grows more competitive with scale, matching full model tuning once models exceed billions of parameters. Soft prompts also improve robustness to domain transfer and enable efficient prompt ensembling.

Based on: The Power of Scale for Parameter-Efficient Prompt Tuning · Conference on Empirical Methods in Natural Language Processing

HighlightarXiv

RoFormer: Enhanced Transformer with Rotary Position Embedding

Introduces Rotary Position Embedding (RoPE), encoding absolute position via a rotation matrix and adding relative-position dependency in self-attention.

The paper investigates how to integrate positional information into transformer-based language models and proposes Rotary Position Embedding (RoPE). RoPE encodes absolute position with a rotation matrix while incorporating explicit relative-position dependency into self-attention. It offers flexibility in sequence length, decaying inter-token dependency with distance, and compatibility with linear self-attention. Evaluated as RoFormer on long-text classification benchmarks, it consistently outperforms alternatives and is supported by theoretical analysis.

Based on: RoFormer: Enhanced Transformer with Rotary Position Embedding · Neurocomputing

HighlightarXiv

Measuring Mathematical Problem Solving With the MATH Dataset

Introduces MATH, a dataset of 12,500 competition math problems with step-by-step solutions for measuring and teaching mathematical reasoning in ML models.

Mathematical problem solving remains difficult for computers. The authors introduce MATH, a dataset of 12,500 challenging competition math problems, each with a full step-by-step solution usable to teach models to generate derivations and explanations. They also release a large auxiliary pretraining dataset covering math fundamentals. Despite some gains, accuracy stays low even with enormous Transformers, and the authors argue that simply scaling model size and compute is impractical, so new algorithmic advances are likely needed.

Based on: Measuring Mathematical Problem Solving With the MATH Dataset · NeurIPS Datasets and Benchmarks

HighlightarXiv

Prefix-Tuning: Optimizing Continuous Prompts for Generation

Proposes prefix-tuning, a lightweight alternative to fine-tuning that freezes the language model and optimizes continuous task-specific prefix vectors.

Fine-tuning adapts large pretrained language models but modifies all parameters, requiring a full model copy per task. Prefix-tuning instead keeps the model frozen and optimizes a small sequence of continuous, task-specific vectors, the prefix, that later tokens attend to as virtual tokens. Applied to GPT-2 for table-to-text and BART for summarization, it learns only 0.1% of the parameters yet matches full fine-tuning with full data, beats it in low-data settings, and extrapolates better to unseen topics.

Based on: Prefix-Tuning: Optimizing Continuous Prompts for Generation · Annual Meeting of the Association for Computational Linguistics

HighlightarXiv

Longformer: The Long-Document Transformer

Introduces Longformer, a transformer whose attention scales linearly with sequence length to process documents of thousands of tokens.

Standard transformers cannot process long sequences because self-attention scales quadratically with length. Longformer introduces attention that scales linearly, combining local windowed attention with task-motivated global attention as a drop-in replacement for self-attention. It reaches state-of-the-art results on character-level language modeling and, when pretrained and finetuned, consistently outperforms RoBERTa on long-document tasks, with new records on WikiHop and TriviaQA. A Longformer-Encoder-Decoder variant supports generative tasks like arXiv summarization.

Based on: Longformer: The Long-Document Transformer · arXiv.org

HighlightarXiv

On the Properties of Neural Machine Translation: Encoder–Decoder Approaches

Analyzes encoder-decoder neural machine translation models, showing performance drops with longer sentences and more unknown words.

Neural machine translation is a then-new approach to statistical machine translation built purely from neural networks, using an encoder that maps a variable-length sentence to a fixed-length representation and a decoder that generates the translation. The paper analyzes two models: an RNN Encoder-Decoder and a newly proposed gated recursive convolutional neural network. Quality is good for short sentences without unknown words but degrades rapidly as sentence length and unknown words grow; the gated model learns grammatical structure automatically.

Based on: On the Properties of Neural Machine Translation: Encoder–Decoder Approaches · SSST@EMNLP

HighlightarXiv

ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

ALBERT introduces two parameter-reduction techniques and an inter-sentence coherence loss to scale BERT pretraining with less memory and faster training.

Scaling up model size in language representation pretraining tends to improve downstream performance but eventually runs into GPU/TPU memory limits and longer training times. ALBERT proposes two parameter-reduction techniques that cut memory use and speed up BERT training, allowing it to scale far better than the original. It also adds a self-supervised loss modeling inter-sentence coherence, which helps tasks with multi-sentence inputs. The best model sets new state-of-the-art results on GLUE, RACE, and SQuAD while using fewer parameters than BERT-large.

Based on: ALBERT: A Lite BERT for Self-supervised Learning of Language Representations · International Conference on Learning Representations

HighlightarXiv

Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling

Empirically compares gated recurrent units (LSTM, GRU) against traditional tanh units in RNNs on polyphonic music and speech signal modeling tasks.

This paper compares different types of recurrent units in recurrent neural networks, focusing on sophisticated units with gating mechanisms such as the long short-term memory (LSTM) unit and the recently proposed gated recurrent unit (GRU). The units are evaluated on polyphonic music modeling and speech signal modeling tasks. Experiments show that the advanced gated units outperform traditional recurrent units such as tanh units, and that the GRU is comparable to the LSTM.

Based on: Empirical Evaluation of Gated Recurrent Neural Networks on Sequence Modeling · arXiv.org

HighlightarXiv

LoRA: Low-Rank Adaptation of Large Language Models

Proposes LoRA, which freezes pre-trained weights and injects trainable low-rank decomposition matrices into Transformer layers for efficient adaptation.

Full fine-tuning of large pre-trained language models becomes impractical at scale — deploying independent fine-tuned instances of GPT-3 175B is prohibitively expensive. LoRA freezes pre-trained weights and injects trainable rank decomposition matrices into each Transformer layer, cutting trainable parameters for downstream tasks by 10,000x and GPU memory by 3x versus fine-tuning GPT-3 with Adam. LoRA matches or beats fine-tuning quality on RoBERTa, DeBERTa, GPT-2, and GPT-3, with higher training throughput and, unlike adapters, no added inference latency.

Based on: LoRA: Low-Rank Adaptation of Large Language Models · International Conference on Learning Representations

HighlightarXiv

LLaMA: Open and Efficient Foundation Language Models

Introduces LLaMA, foundation language models (7B-65B) trained solely on publicly available data, with LLaMA-13B outperforming GPT-3 on most benchmarks.

LLaMA is a collection of foundation language models ranging from 7B to 65B parameters, trained on trillions of tokens. The work shows that state-of-the-art models can be trained using publicly available datasets exclusively, without proprietary or inaccessible data. LLaMA-13B outperforms the 175B GPT-3 on most benchmarks, LLaMA-65B is competitive with Chinchilla-70B and PaLM-540B, and all models are released to the research community.

Based on: LLaMA: Open and Efficient Foundation Language Models · arXiv.org

HighlightarXiv

Sequence to Sequence Learning with Neural Networks

Presents an end-to-end sequence-to-sequence learning approach using multilayered LSTMs to encode inputs to a fixed vector and decode target sequences.

Deep neural networks perform well on difficult tasks given large labeled training sets but cannot map sequences to sequences. This paper presents a general end-to-end sequence learning approach that uses a multilayered LSTM to encode the input sequence into a fixed-dimensional vector and a second deep LSTM to decode the target sequence. On WMT-14 English-to-French translation it reaches 34.8 BLEU versus 33.3 for a phrase-based SMT system, and 36.5 when reranking that system's 1000 hypotheses; reversing source word order markedly improved performance.

Based on: Sequence to Sequence Learning with Neural Networks · Neural Information Processing Systems