Probing the Origins of Reasoning Performance: Representational Quality for Mathematical Problem-Solving in RL vs. SFT Fine-Tuned Models
Quick Answer
This paper shows that Reinforcement learning (RL) models outperform supervised fine-tuned (SFT) models in mathematical reasoning due to superior internal representations.
Quick Take
Linear probes indicate RL models have more structured representations, while hierarchical architecture shows deeper layers are more critical. Token allocation variability suggests adaptive compute allocation is influenced by training pipelines rather than model type alone.
Key Points
- RL models achieve higher accuracy in predicting answer correctness than SFT models.
- Mean ablation studies show RL models develop a hierarchical architecture.
- Deeper layers in RL models become progressively more critical for performance.
- Token-count variability indicates adaptive compute allocation across different models.
- Variability in token allocation suggests influence from overall training pipelines.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large reasoning models trained via reinforcement learning (RL) have been increasingly shown to outperform their supervised fine-tuned (SFT) counterparts on mathematical reasoning tasks; Yet the mechanistic basis for this advantage remains unclear. We therefore ask, what internal representational differences enable RL models' superior performance? Our work presents two converging lines of evidence: First, linear probes trained on layer-wise hidden states reveal that RL models tend to achieve higher accuracy in predicting answer correctness compared to SFT models, indicating more linearly separable and structured representations. Second, mean ablation studies show that RL models develop a hierarchical architecture where deeper layers become progressively more critical, whereas SFT models distribute importance uniformly across layers. Together, these findings demonstrate that RL training fundamentally restructures how models represent and process reasoning problems. Finally, we analyze token-count variability under repeated sampling across problems to assess adaptive compute allocation. While we observe higher variability in some RL-tuned models than in their SFT counterparts, we see strong consistency in others, suggesting that token allocation may depend more on the overall training pipeline than on RL versus SFT alone. We believe this token-allocation variability reveals the spread of plausible on-policy reasoning, highlighting which models exhibit stable policies versus those that are under-determined, potentially non-identifiable solution behaviour.
| Comments: | Second Workshop on XAI4Science, AAAI 2026 |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.26119 [cs.AI] |
| (or arXiv:2607.26119v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26119 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Aishwarya Balwani [view email]
[v1]
Tue, 28 Jul 2026 17:42:42 UTC (4,100 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.