HighlightarXiv

CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph

A framework for elevating web-scale corpus construction to structured knowledge organization.

The paper presents Cortex, a three-layer heterogeneous structure (Ontological Corpus Graph) for organizing high-quality corpora. It refines content, evolves ontologies, and enables cross-domain alignment. Comprehensive experiments validate its effectiveness in quality refinement, domain organization, and data synthesis.

Based on: CORTEX: High-Quality Cross-Domain Organization of Web-Scale Corpora through Ontological Corpus Graph · arXiv

HighlightData & Knowledge Engineering

Semantic constraint validation in knowledge representation for the semantic web: A survey, taxonomy and research challenges

A survey, taxonomy and research challenges on semantic constraint validation.

This paper presents a survey and taxonomy of semantic constraint validation techniques for knowledge representation on the Semantic Web. It identifies research challenges and gaps in current approaches. The authors aim to provide a comprehensive overview of existing methods and their limitations.

Based on: Semantic constraint validation in knowledge representation for the semantic web: A survey, taxonomy and research challenges · Data & Knowledge Engineering

HighlightComputer Science Review

From vectors to knowledge graphs: A comprehensive analysis of modern retrieval-augmented generation architectures

A study on modern retrieval-augmented generation architectures.

The paper analyzes the evolution of retrieval-augmented generation (RAG) models from vector-based representations to knowledge graph-based ones. It provides a comprehensive overview of the current state-of-the-art in RAG architectures and their applications. The authors discuss the benefits and limitations of using knowledge graphs in RAG models.

Based on: From vectors to knowledge graphs: A comprehensive analysis of modern retrieval-augmented generation architectures · Computer Science Review

HighlightarXiv

Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use

A study on the performance of a knowledge-graph tool use recipe with large language models.

The authors test a standard recipe for using knowledge graphs with large language models, observing a 'peak-then-collapse' pattern in performance. They identify four recurring failure modes and argue that interface feedback is a key difference from other tools. The study also explores the effect of self-distillation as a mitigation strategy.

Based on: Peak-Then-Collapse and the Four Interface Channels of Knowledge-Graph Tool Use · arXiv

HighlightarXiv

GRASP: Plan-Guided Graph Retrieval with Adaptive Fusion and Reranking on Semi-Structured Knowledge Bases

A plan-guided graph retrieval method for semi-structured knowledge bases.

The paper proposes GRASP, a plan-guided graph retrieval system that uses adaptive fusion and reranking to improve the accuracy of retrieving relevant information from semi-structured knowledge bases. The approach combines a planning module with a retrieval module to adaptively select relevant subgraphs and fuse their representations. This allows for more accurate and efficient retrieval of relevant information.

Based on: GRASP: Plan-Guided Graph Retrieval with Adaptive Fusion and Reranking on Semi-Structured Knowledge Bases · arXiv

Highlightopenalex.org

When to Trust: A Causality-Aware Calibration Framework for Accurate Knowledge Graph Retrieval-Augmented Generation

A framework that calibrates knowledge graph retrieval-augmented generation models.

The paper proposes Ca2KG, a causality-aware calibration framework for KG-RAG. It integrates counterfactual prompting and a panel-based re-scoring mechanism to improve calibration while maintaining predictive accuracy. Experiments on two QA datasets demonstrate the effectiveness of Ca2KG.

Based on: When to Trust: A Causality-Aware Calibration Framework for Accurate Knowledge Graph Retrieval-Augmented Generation

HighlightarXiv

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering

A schema-aware cumulative process reward model for knowledge graph question answering.

The paper proposes SCPRM, a schema-aware cumulative process reward model for knowledge graph question answering. It aims to improve the performance of KGQA systems by incorporating schema information and cumulative rewards. The authors evaluate SCPRM on several benchmark datasets and demonstrate its effectiveness compared to existing methods.

Based on: SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering · arXiv

HighlightarXiv

Covering the Unseen: Information Demand Coverage Optimization for Retrieval-Augmented Generation

Paper on optimizing information demand coverage in retrieval-augmented generation.

This paper proposes a method to optimize information demand coverage in retrieval-augmented generation (RAG) models. The approach aims to improve the ability of RAG models to retrieve relevant information from external sources. The authors evaluate their method on several benchmarks and demonstrate its effectiveness in improving the performance of RAG models.

Based on: Covering the Unseen: Information Demand Coverage Optimization for Retrieval-Augmented Generation · arXiv

HighlightarXiv

Half a Link can Be Enough to Predict a Whole Link: Understanding Generalization in Knowledge Graph Foundation Models

A research paper exploring generalization in knowledge graph foundation models.

The authors investigate the ability of knowledge graph foundation models to generalize from partial links. They examine whether providing half a link is sufficient for accurate predictions. The study aims to understand how these models can be improved for more efficient and effective knowledge retrieval.

Based on: Half a Link can Be Enough to Predict a Whole Link: Understanding Generalization in Knowledge Graph Foundation Models · arXiv

HighlightarXiv

RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation

Paper proposing a method for multi-hop knowledge graph question answering using recurrent soft-flow and decoupled large language model generation.

The paper introduces RSF-GLLM, a framework that addresses the semantic gap in multi-hop knowledge graph question answering. It combines recurrent soft-flow with decoupled large language model generation to improve performance. The method is evaluated on several benchmarks and shows promising results.

Based on: RSF-GLLM: Bridging the Semantic Gap in Multi-Hop Knowledge Graph QA via Recurrent Soft-Flow and Decoupled LLM Generation · arXiv

HighlightarXiv

BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration

A paper proposing a method for bootstrapping e-commerce attribute taxonomies using human-AI collaboration.

The authors present BEATS, an approach to iteratively refine e-commerce attribute taxonomies through human-AI collaboration. This method aims to improve search performance by leveraging AI-driven suggestions and human feedback. The proposed framework is designed to be applicable in various e-commerce scenarios.

Based on: BEATS: Bootstrapping E-commerce Attribute Taxonomies for Search through Iterative Human-AI Collaboration · arXiv

HighlightarXiv

Beyond Probabilistic Similarity: Structural, Temporal, and Causal Limitations of Retrieval-Augmented Generation in the Legal Domain

A paper discussing limitations of retrieval-augmented generation in the legal domain.

The authors examine structural, temporal, and causal limitations of retrieval-augmented generation (RAG) in the legal domain. They argue that RAG models have inherent biases and limitations when dealing with complex legal concepts and relationships. The paper highlights the need for more robust and accurate methods to handle legal knowledge graphs.

Based on: Beyond Probabilistic Similarity: Structural, Temporal, and Causal Limitations of Retrieval-Augmented Generation in the Legal Domain · arXiv