Right or Wrong, Models Comply: Directional Blindness in LLM Moral Judgment
Quick Answer
This study introduces Compliance Asymmetry (A = BCR/HCR) to evaluate LLMs' responses to nudges, revealing that models exhibit directional blindness in moral judgments, following helpful and harmful nudges equally (A = 1.04), while favoring helpful nudges in factual contexts (A = 1.58).
Quick Take
The findings suggest a need for alignment strategies focusing on directionally calibrated updates.
Key Points
- Compliance Asymmetry measures ' responses to helpful vs. harmful nudges.
- Models show equal compliance to moral nudges (A = 1.04) but favor helpful nudges in factual contexts (A = 1.58).
- Chain-of-thought prompting amplifies compliance for both helpful and harmful nudges.
- Identity-based prompting suppresses compliance for both types of nudges equally.
- Direction-blind moral compliance is identified as a failure mode in current LLMs.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2606. 14037v1 Announce Type: new Abstract: As language models take integrated roles across many domains, the response of to user pushback becomes a critical alignment property. Yet many existing evaluations treat compliance as unidirectional, measuring whether models resist pressure but not whether they resist it selectively. We introduce Compliance Asymmetry (A = BCR/HCR), a bidirectional diagnostic that compares beneficial output change under helpful nudges with harmful change under misleading nudges.
Across 9 models and 972,000 nudge-condition responses, we find that this selectivity differs in factual and moral judgments: models follow helpful nudges more than harmful ones on factual questions (A = 1. 58), but follow both directions at nearly identical rates on moral questions (A = 1. 04). This phenomenon persists across model families, capability levels, and nudging types.
Interestingly, we also find that chain-of-thought prompting amplifies helpful and harmful compliance together, while identity-based prompting suppresses both by nearly identical margins. These results identify direction-blind moral compliance as a distinct failure mode in current LLMs and suggest that alignment should target directionally calibrated updating rather than lower compliance alone.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.