Counterfactual Graph for Multi-Agent LLM Calibration

arXiv cs.CL·Jiatan Huang, Mingchen Li, Ziming Li, Sunjae Kwon, Hong Yu, Chuxu Zhang

4h ago

·~1 min·6/1/2026·en·0

Quick Take

The CAGE-CAL framework enhances multi-agent LLM reliability by calibrating confidence based on counterfactual graphs, improving reliability discrimination with competitive Expected Calibration Error (ECE) across five benchmarks. It addresses the pitfalls of false consensus in agent communication, leading to better topology selection than fixed-topology strategies.

Key Points

CAGE-CAL compares post-communication and no-communication agent graphs.
It captures pairwise failure correlations and group-level dependencies.
The framework improves reliability discrimination across five benchmarks.
CAGE-CAL enhances topology selection beyond fixed-topology strategies.
It addresses over-confidence issues in multi-agent systems.

Article Excerpt

From source RSS / original summary

arXiv:2605. 30653v1 Announce Type: new Abstract: Multi-agent LLM systems often treat agreement as evidence: when many agents in a panel give the same answer, that answer is assumed to be more reliable. We show that this assumption can fail after agents communicate. Communication can induce correlated failures and false consensus, so the same vote share may reflect reliable agreement in one topology but over-confidence in another.

We propose CAGE-CAL, a counterfactual agent-graph calibration framework for multi-agent LLMs. For each query, CAGE-CAL compares an observed post-communication agent graph with a matched counterfactual no-communication graph, capturing both pairwise failure correlations and group-level dependencies. Rather than simply counting how many agents agree, CAGE-CAL estimates the counterfactual shift between observed and no-communication dependence, and calibrates confidence accordingly.

Across five benchmarks, CAGE-CAL improves reliability discrimination with competitive ECE, and its calibrated confidence further improves topology selection over the best fixed-topology strategy.

Reader Mode unavailable (could not extract clean content).

Read on arxiv.org

Want this in your inbox every morning?

Daily brief at your local 8am — bilingual EN/中文, free.

Subscribe — it's free

More from arXiv cs.CL

See more →

arXiv cs.CL·Leyao Wang, Yanan He, Peng Chen, Asaf Yehudai, Yixin Liu, Rex Ying, Michal Shmueli-Scheuer, Arman Cohan

1w ago

FeaturedOriginal

Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?

AI Summary

The REFLECT benchmark reveals that current LLM judges are unreliable, achieving below 55% accuracy in evaluating reasoning and evidence use, highlighting the need for improved evaluation methods for deep research agents.

#LLM #Agent #Inference #Policy