The Formalism Trap: Are LLM-as-a-Judge Evaluators Blinded by Consensus Mimicry under Social Load?
Quick Answer
The study introduces the Agentic Formalism Trap and Evaluative Dissonance Index ($D_E$), revealing how LLM-as-a-Judge systems conflate proceduralism with semantic truth under adversarial conditions.
Quick Take
Analyzing 22,500 trajectories across three domains, the authors identify semantic hallucination patterns and demonstrate that evaluation vulnerabilities are domain-agnostic, necessitating architecture-specific vigilance to mitigate systemic divergence.
Key Points
- Introduces the Evaluative Dissonance Index ($D_E$) for -as-a-Judge systems.
- Analyzed 22,500 trajectories across GAIA, , and Multi-Challenge domains.
- Identified semantic hallucination maneuvers with a validation p-value of $p < 10^{-120}$.
- Logistic meta-evaluator achieved ROC-AUC of 0.8779 for syntactic trigger isolation.
- Vulnerability to evaluator capture is universally domain-agnostic with mean ROC-AUC of 0.7482.
Paper Resources
Source Excerpt
We introduce the \textit{Agentic Formalism Trap} and the Evaluative Dissonance Index ($D_E$), quantifying how -as-a-Judge systems conflate structural proceduralism with semantic truth under adversarial load. Analyzing 22,500 trajectories across 3 domains (GAIA, , Multi-Challenge), we extract a semantic taxonomy of hallucination maneuvers, validated via deterministic lexical grounding ($p < 10^{-120}$). A logistic meta-evaluator isolates the exact syntactic triggers of this evaluator
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.