检索增强生成系统鲁棒性评估及提升框架
首发时间:2026-04-03
摘要:检索增强生成(Retrieval-Augmented Generation,RAG)系统通过整合外部知识库,提高了大型语言模型(Large Language Model,LLM)的事实准确性。然而,其鲁棒性不足可能会导致LLM输出不真实的幻觉信息。针对RAG鲁棒性不足问题,本文从评估与增强两个维度展开研究,提出了一套完整的RAG系统可靠性提升方案。在评估层面,提出RAISP(Robustness and Availability Assessment via Internal Sensitive Knowledge Profiling)框架,通过基于梯度的后缀优化策略在可控环境下模拟检索偏差,量化分析RAG系统面对轻微干扰时的脆弱性;在增强层面,提出PGDRAG(Purification and Geometric Denoising RAG)框架,构建"查询净化-几何去噪-分级验证"三阶段流水线。在多个数据集上的实验结果表明,RAISP评估框架能够有效揭示系统薄弱环节,PGDRAG防御框架显著降低了对抗干扰下的幻觉生成率,为RAG系统在高风险领域的可靠部署提供了技术支撑。
For information in English, please click here
Frameworks for Robustness Evaluation and Enhancement of Retrieval-Augmented Generation
Abstract:Retrieval-Augmented Generation (RAG) systems enhance the factual accuracy of Large Language Model (LLM) by integrating external knowledge bases. However, insufficient robustness in these systems can lead to the generation of non-factual hallucinated information by the LLM. To address the robustness deficiencies in RAG, this research explores the problem from the dual dimensions of evaluation and enhancement, proposing a comprehensive framework for improving RAG system reliability. At the evaluation level, we introduce the RAISP (Robustness and Availability Assessment via Internal Sensitive Knowledge Profiling) framework. This framework simulates retrieval biases in controlled settings via a gradient-based suffix optimization strategy, enabling a quantitative analysis of a RAG system\'s vulnerability to minor perturbations. At the enhancement level, we propose the PGDRAG (Purification and Geometric Denoising RAG) framework, which constructs a three-stage pipeline comprising "Query Purification, Geometric Denoising, and Hierarchical Verification." Experimental results on multiple datasets demonstrate that the RAISP evaluation framework can effectively expose systemic vulnerabilities, while the PGDRAG defense framework significantly reduces the hallucination rate under adversarial interference. This work provides technical support for the reliable deployment of RAG systems in high-stakes domains.
Keywords: Large Language Model Retrieval-Augmented Generation Robustness Enhancement Hallucination
基金:
引用

No.****
动态公开评议
共计0人参与
勘误表
检索增强生成系统鲁棒性评估及提升框架
评论
全部评论