HighlightAAAI Conference on Artificial Intelligence

Learning Entity and Relation Embeddings for Knowledge Graph Completion

Proposes TransR, a knowledge graph embedding model that learns entity and relation embeddings in separate spaces for link prediction.

This paper tackles knowledge graph completion via graph embeddings. Prior models like TransE and TransH treat a relation as a translation from head to tail entity but place entities and relations in one space. Since an entity has multiple aspects that different relations emphasize, the authors propose TransR to build entity and relation embeddings in separate spaces, projecting entities into a relation-specific space before translating. On link prediction, triple classification, and relational fact extraction, TransR gives consistent gains over TransE and TransH.

Based on: Learning Entity and Relation Embeddings for Knowledge Graph Completion · AAAI Conference on Artificial Intelligence

HighlightarXiv

CodeBERT: A Pre-Trained Model for Programming and Natural Languages

Presents CodeBERT, a bimodal Transformer pre-trained on natural language and programming language for code search and documentation.

CodeBERT is a bimodal pre-trained model for programming and natural language that learns general-purpose representations supporting tasks like code search and code documentation generation. Built on a Transformer architecture, it is trained with a hybrid objective including replaced token detection, letting it use both bimodal NL-PL pairs and unimodal data. Fine-tuned, CodeBERT achieves state-of-the-art results on natural language code search and code documentation generation, and it outperforms prior models on a zero-shot NL-PL probing dataset.

Based on: CodeBERT: A Pre-Trained Model for Programming and Natural Languages · Findings

HighlightarXiv

Finetuned Language Models Are Zero-Shot Learners

This paper shows instruction tuning, finetuning a large language model on many tasks phrased as instructions, greatly improves zero-shot generalization.

This paper studies instruction tuning as a simple way to improve the zero-shot abilities of large language models. A 137B-parameter pretrained model is finetuned on over 60 NLP tasks expressed via natural language instruction templates, producing a model called FLAN. On held-out task types, FLAN clearly beats its unmodified counterpart and outperforms zero-shot 175B GPT-3 on 20 of the 25 evaluated tasks, even topping few-shot GPT-3 on benchmarks like ANLI, RTE, and BoolQ. Ablations show the number of finetuning tasks, model scale, and instruction phrasing are all crucial.

Based on: Finetuned Language Models Are Zero-Shot Learners · International Conference on Learning Representations

HighlightarXiv

Survey of Hallucination in Natural Language Generation

Surveys hallucination in natural language generation: metrics, mitigation, and progress across summarization, dialogue, QA, and machine translation.

Natural language generation has advanced rapidly with sequence-to-sequence and Transformer-based language models, yielding more fluent output for tasks like summarization, dialogue, and data-to-text. However, such models often hallucinate unintended text, degrading performance and user trust. This survey comprehensively reviews hallucination in NLG in two parts: a general overview of metrics, mitigation methods, and future directions; and task-specific progress in abstractive summarization, dialogue generation, generative QA, data-to-text, and machine translation.

Based on: Survey of Hallucination in Natural Language Generation · ACM Computing Surveys

HighlightarXiv

Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Presents GNMT, Google's deep LSTM neural machine translation system with attention, wordpieces, and low-precision inference for production use.

Neural machine translation (NMT) enables end-to-end translation but is costly to train and run and handles rare words poorly. GNMT uses a deep LSTM with 8 encoder and 8 decoder layers plus attention and residual connections, wiring the decoder's bottom layer to the encoder's top layer for parallelism. It uses low-precision inference for speed, sub-word 'wordpieces' for rare words, and length-normalized beam search with a coverage penalty. On WMT'14 benchmarks it rivals the state of the art, and human evaluation shows 60% fewer errors than Google's phrase-based system.

Based on: Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation · arXiv.org

HighlightNeural Computation

A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures

A review of LSTM cells and their variants, categorizing LSTM network architectures and surveying applications and future research directions.

Recurrent neural networks are widely used for sequential data such as text, audio, and video, but RNNs built from sigma or tanh cells cannot learn relevant information when the input gap is large. By introducing gate functions into the cell, the long short-term memory (LSTM) handles long-term dependencies well, and most notable RNN results have since relied on it. This review examines the LSTM cell and its variants, divides LSTM networks into LSTM-dominated and integrated categories, discusses applications, and outlines future research directions.

Based on: A Review of Recurrent Neural Networks: LSTM Cells and Network Architectures · Neural Computation

HighlightarXiv

LSTM: A Search Space Odyssey

Presents the first large-scale analysis of eight LSTM variants on three tasks, finding the forget gate and output activation most critical.

Since the LSTM's 1995 inception, many variants have become state-of-the-art, raising interest in which components matter. This paper reports the first large-scale comparison of eight LSTM variants on speech recognition, handwriting recognition, and polyphonic music modeling. Hyperparameters were tuned per task by random search and ranked by functional ANOVA, over 5400 runs (~15 years of CPU time). No variant significantly beats the standard LSTM; the forget gate and output activation are its most critical components, and its hyperparameters are largely independent.

Based on: LSTM: A Search Space Odyssey · IEEE Transactions on Neural Networks and Learning Systems

HighlightarXiv

Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network

A tutorial that formally derives RNN and LSTM equations from differential equations, justifies unrolling, and proposes a generalized Vanilla LSTM.

LSTM networks are widely covered, but most articles state inference formulas axiomatically, omit training formulas, and present RNN 'unrolling' without justification. This tutorial explains essential RNN and LSTM fundamentals in one document. Drawing on signal processing, it formally derives the canonical RNN from differential equations and proves a statement yielding the unrolling technique. It then transforms the RNN into a Vanilla LSTM through logical arguments, provides all governing equations, and introduces extensions producing the most general LSTM variant to date.

Based on: Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) Network · Physica A: Statistical Mechanics and its Applications

HighlightarXiv

On the Opportunities and Risks of Foundation Models

A comprehensive report characterizing foundation models—their capabilities, technical principles, applications, and societal impact.

This report characterizes 'foundation models'—models like BERT, DALL-E, and GPT-3 trained on broad data at scale and adaptable to many downstream tasks. It surveys their opportunities and risks: capabilities (language, vision, robotics, reasoning), technical principles (architectures, training, data), applications (law, healthcare, education), and societal impact (inequity, misuse, environmental effects). Built on deep and transfer learning, their scale yields emergent capabilities and drives homogenization, whose inherited defects demand caution.

Based on: On the Opportunities and Risks of Foundation Models · arXiv.org

HighlightarXiv

Parameter-Efficient Transfer Learning for NLP

Introduces adapter modules that add few trainable parameters per task, enabling parameter-efficient transfer learning for NLP.

Fine-tuning large pre-trained models is effective for NLP transfer but parameter-inefficient, requiring a full new model per task. The authors propose transfer via adapter modules, which add only a few trainable parameters per task while keeping the original network fixed, yielding compact, extensible models with high parameter sharing. Transferring BERT to 26 text classification tasks, including GLUE, adapters reach within 0.4% of full fine-tuning while adding only 3.6% of parameters per task, versus 100% for fine-tuning.

Based on: Parameter-Efficient Transfer Learning for NLP · International Conference on Machine Learning

HighlightConference on Fairness, Accountability and Transparency

On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜

An examination of ever-larger language models, weighing their risks and recommending cost-aware, well-documented, stakeholder-driven alternatives.

The paper steps back from three years of ever-larger English language models such as BERT, GPT-2/3, and Switch-C, which pushed benchmark state of the art through architecture and sheer size. It asks how big is too big and what risks the technology poses, plus paths to mitigate them. The authors recommend weighing environmental and financial costs first, curating and documenting datasets rather than ingesting everything on the web, running pre-development checks of fit with research goals and stakeholder values, and pursuing directions beyond ever-larger models.

Based on: On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? 🦜 · Conference on Fairness, Accountability and Transparency

HighlightarXiv

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Introduces Mamba, an attention-free selective state space model for linear-time sequence modeling across language, audio, and genomics.

Most foundation models rely on the Transformer's attention module, while subquadratic alternatives like structured state space models (SSMs) have lagged on language. The authors trace this to weak content-based reasoning and let SSM parameters depend on the input, so the model selectively propagates or forgets information per token. A hardware-aware parallel algorithm and a simplified attention- and MLP-free design yield Mamba, which offers fast inference, linear scaling, and state-of-the-art results across modalities; Mamba-3B matches Transformers twice its size.

Based on: Mamba: Linear-Time Sequence Modeling with Selective State Spaces · arXiv.org