Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization
Quick Answer
This paper shows that The Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework enhances sexism detection in NLP by clustering annotators based on labeling behavior, fine-tuning language models for each cluster, and optimizing preferences to maintain diverse perspectives.
Quick Take
Evaluated on the EXIST 2024 dataset, MAP-PO demonstrates that cluster-specific training is essential for accurate annotation reproduction.
Key Points
- MAP-PO retains diverse perspectives in sexism detection, unlike majority vote methods.
- Clusters are formed based on annotator behavior, not demographics.
- Fine-tuning language models per cluster improves annotation accuracy.
- Training with shared team-level signals keeps agents aligned with their clusters.
- Results are consistent across English and Spanish tweets in the EXIST 2024 dataset.
DeepSignal Analysis
What happened
The MAP-PO framework aims to improve sexism detection in NLP by clustering annotators based on their labeling behavior. It fine-tunes language models for each cluster and optimizes preferences to reflect diverse perspectives. Evaluated on the EXIST 2024 dataset, the framework shows that cluster-specific training is crucial for accurate annotation reproduction.
Key evidence
- The framework clusters annotators by their labeling behavior, rather than demographic attributes, to capture diverse perspectives.
- MAP-PO was evaluated on the EXIST 2024 dataset, which includes labeled English and Spanish tweets.
- The study found that without fine-tuning, agents behaved similarly, indicating the necessity of cluster-specific training.
Why it matters
This research highlights the importance of recognizing and maintaining diverse perspectives in NLP systems, particularly in sensitive areas like sexism detection. By addressing the limitations of majority voting in annotation, the MAP-PO framework could lead to more nuanced and accurate models. This approach may influence future NLP methodologies and improve the reliability of automated systems in understanding social issues.
Paper Resources
📖 Reader Mode
~2 min readAbstract:When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards. We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine-tuning the agents behave almost identically, so cluster-specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team-level training signal consistently keeps each agent calibrated to its cluster.
| Comments: | 17 pages, 12 figures, 14 tables. Preprint; under review at EACL 2027 (ACL Rolling Review, August 2026 cycle). Code and data: this https URL |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG) |
| Cite as: | arXiv:2608.04056 [cs.CL] |
| (or arXiv:2608.04056v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04056 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hadi Mohammadi [view email]
[v1]
Tue, 4 Aug 2026 11:35:30 UTC (238 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.