Graph database provider Neo4j states graph technology can prevent AI models and agents from generating hallucinations, fabrications and false answers to user queries.
The company referenced a recent academic paper released on arXiv in June, titled “Reducing Hallucinations in Complex Question Answering using Simple Graph-based Retrieval-Augmented Generation”, which validates its claim.
The paper’s abstract notes the study explores adopting a lightweight graph structure with a simple graph schema to support the RAG (Retrieval-Augmented Generation) subsystem via a dedicated toolset. The research team built an agentic system equipped with diverse vector search and graph query tools, operating on structured data sourced from curated English Wikipedia articles. It evaluated system performance on complex questions from MoNaCo, a challenging Wikipedia QA benchmark for complex query answering tasks.
The researchers put forward a complex question for LLMs: “Can you name all the battles between the Dutch and English in the First, Second and Third Anglo-Dutch Wars, and list the victor of each battle?”
Answering this question demands sophisticated retrieval and reasoning capabilities, including simultaneous multi-entity, multi-hop reasoning and cross-document access. The study points out such complex queries pose major challenges for current state-of-the-art LLM-based systems.
The team tested the question on three LLM setups and analyzed the outcomes:
1. Vector+graph RAG: Adopts a unified vector and graph database with predefined tools to optimize external knowledge base retrieval.
2. Simple vector RAG
3. Zero-shot LLM without RAG enhancement
The paper’s findings confirm integrating a basic graph-based knowledge base and matching tools into standard vector RAG can substantially reduce AI hallucinations, though it cannot eliminate them entirely. The coarse truthfulness score improved from approximately −127 for zero-shot LLMs to −49 for vector+graph RAG.
Moreover, the hybrid approach outperforms standalone vector RAG significantly. When partially correct answers are included in evaluation, vector+graph RAG achieves the highest score among the three test scenarios. Its fine-grained truthfulness score is 80 percent higher than that of pure vector RAG, with factual precision and recall more than double those of the vector-only RAG system.
Overall, the proposed solution boosts both precision and recall while curbing hallucinations, offering a viable path to improve the reliability and credibility of LLM-based QA systems.
Read the original paper for comprehensive technical details of the research.
Beijing Qianxing Jietong Technology Co., Ltd.
Sandy Yang/Global Strategy Director
WhatsApp / WeChat: +86 13426366826
Email: yangyd@qianxingdata.com
Website: www.qianxingdata.com/www.storagesserver.com
Business Focus:
ICT Product Distribution/System Integration & Services/Infrastructure Solutions
With 20+ years of IT distribution experience, we partner with leading global brands to deliver reliable products and professional services.
“Using Technology to Build an Intelligent World”Your Trusted ICT Product Service Provider!