How Narrative Wrapping Affects LLM Refusal: A Cross-Language Benchmark and Defense
Quick Answer
This paper shows that Safety-aligned language models, like Qwen3-1.7B, show vulnerability to harmful requests when framed in narratives, achieving 89.4% success in English and 95.7% in Classical Chinese.
Quick Take
The proposed AXIS method enhances refusal capabilities, outperforming existing models in safety and usability across Qwen3 and GLM-4 benchmarks.
Key Points
- Qwen3-1.7B shows 89.4% attack success in English and 93.0% in modern Chinese.
- Classical Chinese requests reach a 95.7% success rate for harmful requests.
- GUISE benchmark includes harmful and benign request pairs across languages.
- AXIS combines preference optimization with a rotation objective for better refusal.
- AXIS achieves the highest safety and usability scores among compared methods.
DeepSignal Analysis
What happened
Research indicates that safety-aligned language models, such as Qwen3-1.7B, are susceptible to harmful requests when presented within narrative contexts. The attack success rates are notably high, reaching 89.4% in English and 95.7% in Classical Chinese. A new method called AXIS has been proposed to enhance refusal capabilities in these models.
Key evidence
- Qwen3-1.7B exhibits an attack success rate of 89.4% for harmful requests framed in narratives in English.
- In Classical Chinese, the attack success rate for the same model rises to 95.7%, indicating a significant vulnerability.
- The AXIS method combines preference optimization and a rotation objective, achieving the highest safety and usability scores across Qwen3 and GLM-4 benchmarks.
Why it matters
The findings highlight a critical vulnerability in language models that can be exploited through narrative framing, raising concerns about their safety in real-world applications. The high success rates of harmful requests suggest that existing safety measures may be inadequate. The introduction of AXIS aims to address these vulnerabilities, potentially improving the reliability of language models in sensitive contexts.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Safety-aligned language models often refuse a harmful request stated directly but answer the same request inside a role-play or narrative wrapper. We measure this vulnerability across languages and registers: attack success on Qwen3-1.7B is already 89.4% in English and 93.0% in modern Chinese, and reaches 95.7% in Classical Chinese. We build GUISE, a benchmark for systematically studying this vulnerability. It includes parallel requests in English, modern Chinese, and Classical Chinese, matched harmful and benign pairs, wrapper types held out for evaluation, and a stricter criterion that counts warn-then-answer responses as attack successes. Representation analysis shows that language and register move harmful-request representations only slightly away from the model's refusal direction, whereas narrative wrappers move them much farther away. We propose AXIS, which combines preference optimisation with a rotation objective that aligns harmful-request representations with the refusal direction and a commitment objective that trains the model to refuse completely rather than produce a warn-then-answer response. Across Qwen3-1.7B, Qwen3-4B and GLM-4-9B, AXIS achieves the highest combined safety and usability score among the compared methods.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11005 [cs.AI] |
| (or arXiv:2610.11005v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11005 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhankai Ye [view email]
[v1]
Wed, 7 Oct 2026 23:40:25 UTC (381 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.