Beyond Refusal Patterns: Safe-Role Internalization for Robust and Generalizable LLM Safety Alignment
Quick Answer
The SSRFT framework introduces a novel approach to safety alignment in LLMs by internalizing a safe role, enhancing robustness against jailbreak attacks and reducing over-refusal.
Quick Take
Experiments show SSRFT outperforms standard SFT in generalizability and safety, preserving model capabilities while addressing vulnerabilities in various scenarios.
Key Points
- SSRFT reformulates safety alignment as internalizing a predefined safe role.
- Constructs Safe-Role Question-Answer dataset from psychometric questions and limited prompts.
- Demonstrates greater robustness to prefilling attacks compared to standard SFT.
- Reduces over-refusal on benign queries while maintaining general model capabilities.
- Establishes safe-role internalization as a viable alternative to refusal-centric approaches.
DeepSignal Analysis
What happened
The SSRFT framework proposes a new method for safety alignment in large language models (LLMs) by focusing on internalizing a safe role. This approach aims to enhance robustness against jailbreak attacks and mitigate issues of over-refusal. Experimental results indicate that SSRFT outperforms standard supervised fine-tuning (SFT) in terms of generalizability and safety.
Key evidence
- SSRFT constructs a Safe-Role Question-Answer dataset from psychometric questions and limited jailbreak prompts, allowing models to internalize safety-oriented values.
- Experiments show that SSRFT achieves greater robustness to prefilling attacks and better generalization to unseen jailbreak domains compared to standard SFT.
- The framework reduces over-refusal on benign queries while preserving the model's general capabilities, indicating a shift from refusal-centric safety alignment.
Why it matters
The introduction of SSRFT addresses significant vulnerabilities in existing safety alignment methods for LLMs, which often rely on extensive supervision and can be easily manipulated. By internalizing a safe role, SSRFT aims to create models that are not only safer but also more adaptable to various scenarios. This could lead to more reliable AI systems that maintain their capabilities while minimizing harmful outputs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large Language Models (LLMs) have achieved remarkable capabilities but remain vulnerable to jailbreak attacks that elicit harmful or unsafe outputs. Existing safety alignment approaches, including Supervised Fine-Tuning (SFT) and Reinforcement Learning from Human Feedback (RLHF), often require substantial attack-specific supervision and computational resources, while remaining susceptible to shallow safety alignment and over-refusal. To address these challenges, we introduce SSRFT(Supervised Safe-Role Fine-Tuning), the first framework that reformulates safety alignment as the internalization of a predefined safe role. SSRFT constructs a Safe-Role Question-Answer (SRQA) dataset from psychometric questions, limited jailbreak prompts, and a safe-role description. Role-consistent responses are synthesized, validated, and expanded into diverse scenarios, enabling models to internalize safety-oriented values and principles rather than explicit refusal patterns. Experiments across multiple Base and Instruct models show that SSRFT achieves more robust and generalizable safety alignment than standard SFT. SSRFT shows substantially greater robustness to prefilling attacks and better generalization to unseen jailbreak domains, while reducing over-refusal on benign queries and preserving the model's general capabilities. These results establish safe-role internalization as an effective alternative to refusal-centric safety alignment. Warning: This paper contains examples of harmful and toxic language.
| Comments: | 27 pages,7 figures, under review |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2610.07023 [cs.AI] |
| (or arXiv:2610.07023v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07023 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jinghao Pang [view email]
[v1]
Sun, 4 Oct 2026 15:50:42 UTC (1,991 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.