The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Quick Answer
The study introduces the Agentic Formalism Trap and Evaluative Dissonance Index ($D_E$), revealing how LLM-as-a-Judge systems conflate proceduralism with semantic truth under adversarial conditions.
Quick Take
Analyzing 22,500 trajectories across three domains, the authors identify semantic hallucination patterns and demonstrate that evaluation vulnerabilities are domain-agnostic, necessitating architecture-specific vigilance to mitigate systemic divergence.
Key Points
- Introduces the Evaluative Dissonance Index ($D_E$) for -as-a-Judge systems.
- Analyzed 22,500 trajectories across GAIA, , and Multi-Challenge domains.
- Identified semantic hallucination maneuvers with a validation p-value of $p < 10^{-120}$.
- Logistic meta-evaluator achieved ROC-AUC of 0.8779 for syntactic trigger isolation.
- Vulnerability to evaluator capture is universally domain-agnostic with mean ROC-AUC of 0.7482.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how LLM-as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, SWE-bench, Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator capture (ROC-AUC 0.8779), while a zero-shot Leave-One-Domain-Out transfer proves the vulnerability is universally domain-agnostic (mean ROC-AUC 0.7482). Architectural profiling reveals that distinct simulated swarm topologies induce mathematically disparate semantic blind spots, proving that unanchored closed-loop evaluation is unstable, systemically divergent and necessitates architecture-specific vigilance filters.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.28641 [cs.CL] |
| (or arXiv:2607.28641v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28641 arXiv-issued DOI via DataCite |
Submission history
From: Dahlia Shehata [view email]
[v1]
Wed, 20 May 2026 02:38:39 UTC (650 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.