Exploring Agentic Tool-Calling Decisions via Uncertainty-Aligned Reinforcement Learning
Quick Answer
The TRUST framework enhances decision-making in LLM-based agents by integrating uncertainty quantification into reward design, improving tool-use outcomes across benchmarks.
Quick Take
Experimental results indicate a consistent increase in decision quality and agent performance while providing reliable uncertainty estimates during optimization.
Key Points
- TRUST incorporates uncertainty quantification to improve decision-making in agents.
- Existing methods often lead to overconfident mistakes in decisions.
- Experimental results show enhanced decision quality across diverse benchmarks.
- The framework maintains reliable uncertainty estimates during optimization.
- Lightweight key-turn annotations are used for unified post-training.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions. Existing approaches mainly improve these behaviors through inference-time correction or coarse-grained reward signals based on decision outcomes and structured checklists, leaving the uncertainty characteristics of agent decisions underexplored. We observe that decision-oriented reinforcement learning tends to weaken the uncertainty separation between correct and incorrect actions, resulting in overconfident mistakes and weaker exploration signals. Therefore, we propose TRUST, which incorporates uncertainty quantification into reward design as a repulsive force for maintaining uncertainty separation, and labels lightweight key-turn annotations for unified post-training of multi-turn trajectories. Experimental results across diverse tool-use benchmarks show that TRUST consistently enhances both decision quality and agent performance while maintaining more reliable uncertainty estimates during optimization.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.06976 [cs.AI] |
| (or arXiv:2606.06976v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.06976 arXiv-issued DOI via DataCite |
Submission history
From: Yijin Zhou [view email]
[v1]
Fri, 5 Jun 2026 07:08:34 UTC (15,276 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.