HighlightarXiv

Training language models to follow instructions with human feedback

Introduces InstructGPT: GPT-3 fine-tuned on demonstrations and human feedback rankings via RL to align language models with user intent.

Larger language models are not inherently better at following user intent and can produce untruthful, toxic, or unhelpful outputs. The authors align models by fine-tuning GPT-3 on labeler-written demonstrations, then further fine-tuning with reinforcement learning from human feedback using rankings of model outputs, producing InstructGPT. In human evaluations, outputs from the 1.3B-parameter InstructGPT are preferred over the 175B GPT-3 despite 100x fewer parameters, with improved truthfulness, less toxic generation, and minimal regressions on public NLP datasets.

Based on: Training language models to follow instructions with human feedback · Neural Information Processing Systems

HighlightarXiv

Evaluating Large Language Models Trained on Code

Introduces Codex, a GPT model fine-tuned on GitHub code, and HumanEval, a benchmark for functional correctness of programs synthesized from docstrings.

The paper introduces Codex, a GPT language model fine-tuned on public GitHub code, and studies its Python code-writing abilities; a distinct production version powers GitHub Copilot. On HumanEval, a newly released benchmark measuring functional correctness of programs synthesized from docstrings, Codex solves 28.8% of problems versus 0% for GPT-3 and 11.4% for GPT-J, and repeated sampling solves 70.2% with 100 samples per problem. The authors also examine limitations, such as long chains of operations and variable binding, and discuss safety, security, and economic impacts.

Based on: Evaluating Large Language Models Trained on Code · arXiv.org

HighlightarXiv

Language Models are Few-Shot Learners

Trains the 175-billion-parameter GPT-3 and shows strong few-shot task performance from text prompts alone, without fine-tuning.

Prior NLP gains came from pretraining plus fine-tuning needing many labeled examples, unlike humans who learn from a few examples. This paper trains GPT-3, a 175-billion-parameter autoregressive language model 10x larger than prior non-sparse models, evaluated via pure in-context few-shot prompting with no gradient updates. GPT-3 performs strongly on translation, QA, cloze, and reasoning tasks, rivaling some fine-tuned methods, though it struggles on some datasets and can write news articles hard to distinguish from human-written ones.

Based on: Language Models are Few-Shot Learners · Neural Information Processing Systems

HighlightarXiv

Neural Machine Translation by Jointly Learning to Align and Translate

Introduces a soft-attention mechanism letting an encoder-decoder network jointly align and translate without fixed-length bottleneck.

Neural machine translation models typically use an encoder-decoder architecture where an encoder compresses a source sentence into a fixed-length vector from which a decoder generates the translation. The authors conjecture this fixed-length vector is a bottleneck, and propose extending the model to automatically (soft-)search source-sentence parts relevant to each target word, without explicit segmentation. This matches existing state-of-the-art phrase-based system performance on English-to-French, and the learned soft alignments agree well with intuition.

Based on: Neural Machine Translation by Jointly Learning to Align and Translate · International Conference on Learning Representations

HighlightarXiv

RoBERTa: A Robustly Optimized BERT Pretraining Approach

A replication study showing BERT was undertrained, and that a tuned pretraining recipe matches or beats later models.

Language model pretraining yields strong gains but careful comparison across approaches is difficult, since training is expensive, done on private datasets of varying sizes, and sensitive to hyperparameter choices. This paper presents a replication study of BERT pretraining that measures the impact of key hyperparameters and training data size, finding BERT was significantly undertrained. A better-tuned BERT can match or exceed every model published after it, achieving state-of-the-art results on GLUE, RACE, and SQuAD; the authors release their models and code.

Based on: RoBERTa: A Robustly Optimized BERT Pretraining Approach · arXiv.org

HighlightAnnual Meeting of the Association for Computational Linguistics

Bleu: a Method for Automatic Evaluation of Machine Translation

Proposes BLEU, a quick, inexpensive, language-independent automatic method for evaluating machine translation quality.

Human evaluation of machine translation is thorough but slow, costly, and its labor cannot be reused. The authors propose an automatic evaluation method that is fast, inexpensive, and language-independent, correlating highly with human judgments while adding little marginal cost per run. It is presented as an automated substitute for skilled human judges, useful whenever quick or frequent evaluation is needed.

Based on: Bleu: a Method for Automatic Evaluation of Machine Translation · Annual Meeting of the Association for Computational Linguistics

HighlightarXiv

Efficient Estimation of Word Representations in Vector Space

Proposes two new architectures for learning continuous word vector representations efficiently from very large datasets.

The paper proposes two novel architectures for computing continuous vector representations of words from very large datasets. Representation quality is measured on a word similarity task and compared against prior neural-network-based techniques. The authors report large accuracy gains at much lower computational cost, learning high-quality word vectors from a 1.6 billion word dataset in under a day. These vectors also achieve state-of-the-art results on a test set measuring syntactic and semantic word similarities.

Based on: Efficient Estimation of Word Representations in Vector Space · International Conference on Learning Representations

HighlightConference on Empirical Methods in Natural Language Processing

GloVe: Global Vectors for Word Representation

Introduces GloVe, a global log-bilinear regression model unifying matrix factorization and context-window methods for word vectors.

Prior word-vector methods captured semantic and syntactic regularities through vector arithmetic, but why these regularities arose was unclear. The authors make explicit the properties needed for such structure to emerge and propose GloVe, a global log-bilinear regression model combining matrix factorization with local context-window approaches. It trains only on nonzero entries of a word-word co-occurrence matrix. The resulting vectors score 75% on a word analogy task and outperform related models on similarity and named entity recognition.

Based on: GloVe: Global Vectors for Word Representation · Conference on Empirical Methods in Natural Language Processing

HighlightarXiv

BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Devlin et al. present BERT, a bidirectional Transformer pretraining method that set new state-of-the-art results on eleven NLP tasks.

BERT pre-trains deep bidirectional representations by jointly conditioning on left and right context in every layer, unlike prior left-to-right language models. A single pretrained BERT model can be fine-tuned with one extra output layer for many tasks, pushing GLUE to 80.5, MultiNLI accuracy to 86.7%, and SQuAD v1.1 F1 to 93.2 — new state-of-the-art results across eleven NLP benchmarks.

Based on: BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding · North American Chapter of the Association for Computational Linguistics

HighlightarXiv

Attention is All you Need

Vaswani et al. propose the Transformer, an architecture built solely on attention, replacing recurrence and convolution.

The Transformer dispenses with recurrent and convolutional layers entirely, relying only on attention mechanisms. It is more parallelizable and faster to train than prior encoder-decoder models, reaching 28.4 BLEU on WMT 2014 English-to-German and a new state-of-the-art 41.8 BLEU on English-to-French after 3.5 days of training on eight GPUs.

Based on: Attention is All you Need · Neural Information Processing Systems

HighlightarXiv

Increasing the LLM Accuracy for Question Answering: Ontologies to the Rescue!

Paper on improving question answering systems with large language models using ontologies.

This paper presents an approach to improve the accuracy of question answering systems with large language models by leveraging ontologies. The authors propose a method that consists of ontology-based query check and LLM repair, which increases the overall accuracy to 72%. The results provide further evidence that investing knowledge graphs, namely the ontology, provides higher accuracy for LLM-powered question-answering systems.

Based on: Increasing the LLM Accuracy for Question Answering: Ontologies to the Rescue! · arxiv.org

Highlightgit.nextgraph.org

oxigraph

A Rust-based graph database implementing the SPARQL standard.

Oxigraph is a graph database written in Rust that implements the SPARQL standard. It provides a compliant, safe, and fast graph database based on the RocksDB key-value store. Oxigraph also includes utility functions for reading, writing, and processing RDF files.

Based on: oxigraph · git.nextgraph.org