Bayesian Uncertainty Propagation for Agentic RAG Pipelines: A Proof-of-Concept Study on Multi-Hop Question Answering
Quick Answer
This study introduces a Bayesian uncertainty-aware framework for Agentic RAG systems, evaluated on StrategyQA and HotpotQA using GPT-3.5-Turbo and GPT-4.1-Nano.
Quick Take
Results indicate that Bayesian propagation is more effective in HotpotQA, highlighting the need for further validation in industrial applications like Offshore Wind maintenance.
Key Points
- Introduces a Bayesian framework for estimating uncertainty in multi-hop systems.
- Evaluated using GPT-3.5-Turbo and GPT-4.1-Nano on StrategyQA and HotpotQA.
- Bayesian propagation shows better performance on HotpotQA compared to StrategyQA.
- Metrics used include AUROC, AUARC, ECE, and Brier Score for assessment.
- Future validation needed in industrial domains like Offshore Wind maintenance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Trustworthy deployment of Agentic Retrieval-Augmented Generation (RAG) systems requires mechanisms for estimating when multi-stage reasoning pipelines may fail. This paper presents an uncertainty-aware Agentic Retrieval-Augmented Generation (RAG) framework in which planner, evaluator and generator stages produce uncertainty signals derived from semantic divergence and generator self-evaluation. These signals are propagated through a Bayesian Network (BN) to estimate system-level uncertainty and provide node-level indicators of potential failure points across the workflow. The approach is evaluated on StrategyQA and HotpotQA using GPT-3.5-Turbo and GPT-4.1-Nano, with Area Under the Receiver Operating Characteristic Curve (AUROC), Area Under the Accuracy-Rejection Curve (AUARC), Expected Calibration Error (ECE), and Brier Score used to assess discrimination, selective prediction and calibration. Results show that Bayesian propagation is more effective on HotpotQA, where uncertainty accumulates across multi-hop reasoning stages, while StrategyQA exposes limitations caused by miscalibration and unreliable upstream signals. The study positions Bayesian uncertainty propagation as a promising but preliminary mechanism for monitoring Agentic RAG systems, with future validation required in industrial domains such as Offshore Wind (OSW) maintenance decision support.
| Comments: | Submitted for 7th International Conference on Maintenance and Intelligent Asset Management (ICMIAM 2026) |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.00972 [cs.AI] |
| (or arXiv:2607.00972v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.00972 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Koorosh Aslansefat [view email]
[v1]
Wed, 1 Jul 2026 14:08:58 UTC (173 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.