The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
Quick Answer
MARS-Gov introduces a multi-agent framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880.
Quick Take
This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.
Key Points
- MARS-Gov combines legal retrieval and open-set target screening for bias detection.
- Achieved 0.880 F1 score, outperforming top zero-shot LLM by 20.2 points.
- Reduced unnecessary interventions to 2.5%, enhancing efficiency in bias governance.
- Dynamic '10th juror' allows for deliberation beyond fixed bias categories.
- Leave-One-Category-Out evaluation shows 85.1% Correct@1 accuracy.
DeepSignal Analysis
What happened
The MARS-Gov framework has been developed to detect bureaucratic bias in Dutch government documents. It achieved a state-of-the-art F1 score of 0.880, surpassing existing zero-shot LLM detectors by 20.2 points. The framework includes a dynamic '10th juror' that adapts to new biases.
Key evidence
- MARS-Gov achieved an F1 score of 0.880, setting a new state-of-the-art in bias detection for Dutch government documents.
- The model outperformed the strongest zero-shot LLM detector by 20.2 points, which is a 29.8% relative improvement.
- Unnecessary interventions were reduced to just 2.5% through the use of the MARS-Gov framework.
Why it matters
This advancement in bias detection is significant as it addresses the limitations of existing methods, which often fail to adapt to emerging biases. By incorporating a dynamic juror system, MARS-Gov enhances the ability to identify and mitigate bias in legal language, potentially improving the fairness of bureaucratic processes.
What to watch
Future evaluations should focus on the framework's adaptability to various types of biases and its performance in real-world applications. Additionally, the implications of reducing unnecessary interventions on bureaucratic efficiency and fairness warrant further investigation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Presupposing the boundaries of bias is itself a form of bias. We study closed-loop bias governance for Dutch government documents, where a system must detect biased language, ground decisions in legal and contextual evidence, rewrite problematic sentences when intervention is warranted, and verify that the rewrite mitigates harm without distorting meaning. Existing methods face three challenges: (i) discriminative classifiers capture surface regularities but lack normative grounding; (ii) zero-shot LLMs often adopt generic viewpoints and over-flag ambiguous administrative language; and (iii) fixed taxonomies inherit the Closed-World Assumption, missing emerging local targets. We propose MARS-Gov, a standpoint-aware multi-agent framework that combines legal retrieval, open-set target screening, specialized jurors, conservative routing, and rewrite verification. When screening finds an uncovered group, MARS-Gov instantiates a dynamic "10th juror" to deliberate outside the fixed panel. On DGDB, MARS-Gov sets a new SOTA with 0.880 F1, outperforming the strongest zero-shot LLM detector by 20.2 points (29.8% relative) and the best supervised Dutch encoder by 6.8 points, while reducing unnecessary interventions to 2.5%. Leave-One-Category-Out (LOCO) evaluation recovers held-out categories with 85.1% Correct@1 and 93.8% Correct@3.
| Comments: | 19 pages, 6 figures. Accepted at EMNLP 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.11136 [cs.CL] |
| (or arXiv:2610.11136v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11136 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zijun Wang [view email]
[v1]
Thu, 8 Oct 2026 03:05:08 UTC (924 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.