Wait, am I Being Fair? Characterizing Deductive Stereotyping and Mitigating It with Fair-GCG
Quick Answer
This paper identifies deductive stereotyping in large language models (LLMs), where models make biased inferences based on population statistics.
Quick Take
To counteract this, the authors propose Fair-GCG, a framework that enhances fairness-aware reasoning by discovering effective injection phrases, leading to improved performance on fairness benchmarks and real-world tasks.
Key Points
- Deductive stereotyping leads to biased inferences in despite improved reasoning.
- Fair-GCG systematically discovers injection phrases to enhance fairness in reasoning.
- The proposed framework improves performance across multiple fairness benchmarks.
- Fair-GCG generalizes from smaller to larger LLMs, reducing bias in outputs.
- The approach transfers effectively to real-world fairness-sensitive applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Warning: This paper contains several toxic and offensive statements. While reasoning generally improves fairness in recent large language models (LLMs), failures persist. In this work, we identify a failure mode, deductive stereotyping, in which models apply population-level statistical regularities to individual cases, producing logically coherent yet socially biased inferences. We provide a statistical interpretation of this phenomenon. To steer models toward fairness-aware reasoning, we propose a reasoning-time injection framework. We further introduce Fair-GCG to systematically discover effective injection phrases. Injection phrases discovered by Fair-GCG improve performance across multiple fairness benchmarks, generalize from smaller to larger LLMs, improves reasoning-level fairness, reduces bias in open-ended generation, and transfer to real-world fairness-sensitive tasks.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.30989 [cs.CL] |
| (or arXiv:2606.30989v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.30989 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Naihao Deng [view email]
[v1]
Tue, 30 Jun 2026 00:00:42 UTC (1,143 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.