Are Arithmetic Heuristic Neurons Form-Invariant? A Mechanistic Analysis of Symbols, Text, and Code in LLMs
Quick Answer
The study investigates arithmetic heuristic neurons in Llama-3 models, revealing that they are form-invariant across symbolic arithmetic, natural language problems, and Python code.
Quick Take
A shared set of neurons enables high accuracy in arithmetic tasks, with over 97% recovery of incorrect predictions when transferring activations between formats, indicating that failures stem from activation states rather than distinct circuits.
Key Points
- Arithmetic heuristic neurons in Llama-3 models are form-invariant across different formats.
- A compact set of neurons is shared among symbolic arithmetic, natural language, and code.
- Transferring activations recovers over 97% of incorrect predictions for addition and subtraction.
- Failures in arise from activation states, not distinct internal circuits.
- Shared neurons belong to the same heuristic families across formats.
DeepSignal Analysis
What happened
The study examines arithmetic heuristic neurons in Llama-3 models, finding that these neurons are form-invariant across different formats such as symbolic arithmetic, natural language problems, and Python code. The research indicates that a shared set of neurons is responsible for high accuracy in arithmetic tasks, with over 97% recovery of incorrect predictions when transferring activations between formats.
Key evidence
- The study identifies arithmetic heuristic neurons in three Llama-3 models using a two-stage pipeline of attribution and activation patching.
- A compact set of neurons is shared across symbolic arithmetic, natural language problems, and Python code, demonstrating their necessity for arithmetic computation.
- Transferring activations from successful executions to failed ones recovers over 97% of incorrect predictions, suggesting that failures are due to activation states rather than separate circuits.
Why it matters
Understanding the form-invariance of arithmetic heuristic neurons in LLMs could lead to improved model designs that leverage shared circuits for better performance across various problem formats. This insight may also inform future research on the interpretability of neural networks and their underlying mechanisms, potentially enhancing their reliability in critical applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models often succeed on one formulation of a problem while failing on an equivalent formulation. Whether these failures arise from distinct internal circuits or different activation states of a shared circuit remains unknown. Recent mechanistic interpretability studies suggest that arithmetic in LLMs emerges from a "bag of heuristics," encoded by a sparse set of MLP neurons that represent distinct arithmetic strategies. We investigate whether arithmetic heuristic neurons are form-invariant across symbolic arithmetic, natural language word problems, and Python code in three Llama-3 models. In each format, we identify arithmetic heuristic neurons using a two-stage pipeline combining attribution patching and activation patching. A compact set of neurons is shared across all three formats, and targeted interventions show this shared circuit is both necessary and sufficient for late-layer arithmetic computation. Transferring the shared neurons' activations from a successful execution in one format to a failed execution in another recovers most incorrect predictions, exceeding 97% for addition and subtraction, indicating that cross-format failures arise from activation states rather than distinct circuits. Moreover, shared neurons consistently belong to the same heuristic families across formats, demonstrating that arithmetic computation in LLMs is largely form-invariant at the neuron level.
| Comments: | Under Review |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.16693 [cs.CL] |
| (or arXiv:2607.16693v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16693 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tanvir Ahmed Sijan [view email]
[v1]
Sat, 18 Jul 2026 08:02:50 UTC (714 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.