Reference Feature Atlases for Mechanistic Auditing of Language Models
Quick Answer
The proposed reference feature atlas enables efficient auditing of language models by reusing a sparse feature library, achieving controllable outcomes on models like Mistral and Qwen-2.5.
Quick Take
The residual channel effectively identifies and manipulates hidden objectives, outperforming traditional baselines in real-time adjustments.
Key Points
- Reference feature atlas trained on five 7-9B instruction-tuned models.
- Residual channel allows for real-time control of hidden objectives.
- Mistral and Qwen-2.5 models show significant performance improvements.
- Traditional baselines failed to achieve the same level of control.
- Panel-relative political framing metrics can be shifted without affecting out-of-domain controls.
DeepSignal Analysis
What happened
The authors introduce a reference feature atlas, a sparse feature library that allows for efficient auditing of language models like Mistral and Qwen-2.5. This atlas enables the identification and manipulation of hidden objectives through a residual channel, which outperforms traditional methods in real-time adjustments.
Key evidence
- The reference feature atlas is trained on a panel and reused for new targets, requiring only a linear decoder for fitting.
- In tests, the residual channel successfully controls hidden objectives in Mistral and Qwen-2.5, while traditional baselines fail to achieve similar results.
- The residual channel also identifies a political-framing cluster in Qwen-2.5, allowing for adjustments in framing metrics without affecting unrelated controls.
Why it matters
This development could significantly enhance the auditing process for language models, allowing for more precise control over their outputs. By reusing a trained feature library, the approach reduces the need for extensive retraining, potentially saving time and resources. The ability to manipulate hidden objectives in real-time could lead to more transparent and accountable AI systems, addressing concerns about bias and unintended consequences in language model outputs.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal.
We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives injected into both targets, the residual channel makes the planted mechanism perfectly controllable at runtime while matched controls stay unaffected and recovers the planted objective as the top-ranked latent across both targets; on Mistral, where the per-target SAE and pairwise crosscoder baselines are retrained for a head-to-head benchmark, both baselines fail to do so. On Qwen-2.5, the same channel additionally reveals a panel-relative political-framing cluster; steering it shifts the audited framing metrics while out-of-domain controls remain unchanged.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.22570 [cs.AI] |
| (or arXiv:2607.22570v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22570 arXiv-issued DOI via DataCite |
Submission history
From: Tong Che [view email]
[v1]
Mon, 1 Jun 2026 17:20:16 UTC (598 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.