The Persona Hierarchy Model: Understanding Contextual Generalization in Fine-Tuning LLMs
Quick Answer
This paper shows that The Persona Hierarchy Model explains how fine-tuning LLMs like Qwen3-4B influences contextual generalization, showing a strong correlation (Pearson's r = 0.72) between training context persona similarity and generalization narrowness.
Quick Take
Implementing persona-preserving regularization (PPR) significantly reduces reward hacking while maintaining accuracy, suggesting new avenues for better alignment in .
Key Points
- Fine-tuning modifies shared personas for broader generalization across contexts.
- Generalization narrowness positively correlates with persona similarity (Pearson's r = 0.72).
- Prior fine-tuning under default contexts enhances future generalization.
- PPR reduces reward hacking from 42-55% to 0.2% while retaining accuracy.
- Results support the Persona Hierarchy Model for contextual generalization.
DeepSignal Analysis
What happened
The Persona Hierarchy Model was proposed to explain how fine-tuning large language models (LLMs) influences their ability to generalize across different contexts. The model indicates that a shared default persona can affect behavior, with fine-tuning that alters this persona leading to broader generalization. Persona-preserving regularization (PPR) was shown to significantly reduce reward hacking while maintaining accuracy.
Key evidence
- The study analyzed 120 fine-tuned models across four behaviors and 15 training contexts, finding a Pearson's correlation of r = 0.72 between persona similarity and generalization narrowness for Qwen3-4B.
- Prior fine-tuning under a default context was found to enhance generalization in subsequent training under different contexts.
- Implementing persona-preserving regularization (PPR) reduced reward hacking from 42-55% to at most 0.2% while retaining accuracy gains.
Why it matters
Understanding the Persona Hierarchy Model provides insights into how LLMs can be fine-tuned for better performance across various contexts. The findings suggest that aligning responses with a default persona can enhance generalization, which is crucial for developing more robust AI systems. Additionally, the reduction of reward hacking through PPR indicates potential for improved alignment in LLMs, addressing concerns about unintended behaviors.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Language models are routinely fine-tuned under a fixed context, such as a generic system prompt, persona or domain-specific instruction, yet the learned behavior sometimes stays confined to that context and sometimes broadly generalizes to unseen contexts. We propose the Persona Hierarchy Model to explain this: a shared default persona influences behavior across contexts. Under this model, fine-tuning that modifies the shared persona promotes broader transfer, whereas changes to local personas remain more context-specific. Across 120 fine-tuned models spanning four behaviors and 15 training contexts, generalization narrowness positively correlates with the similarity between the training context's persona and the default persona (Pearson's r = 0.72 for Qwen3-4B). Prior fine-tuning under the default context can broaden generalization in subsequent training under other contexts. Aligning contextual responses with default-persona responses produces stronger effects. Finally, we propose persona-preserving regularization (PPR) to confine undesired contextual generalization. In RL, PPR cuts reward hacking from 42-55% to at most 0.2% under every evaluated prompt while retaining accuracy gains. These results support the Persona Hierarchy Model as an explanation for contextual generalization and can motivate future controls on unintended generalization for better alignment of LLMs.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09384 [cs.CL] |
| (or arXiv:2610.09384v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09384 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiachen Zhao [view email]
[v1]
Wed, 7 Oct 2026 03:40:44 UTC (1,786 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.