Sentence-Level Contextual Entrainment in Large Language Models
Quick Answer
This study reveals that large language models (LLMs) exhibit sentence-level contextual entrainment, where sentences in prompts can significantly boost token probabilities during inference.
Quick Take
Analyzing 26 , it was found that this effect diminishes with model size and can be mitigated by disabling 2-4% of attention heads without degrading performance.
Key Points
- Sentence-level contextual entrainment boosts token probabilities during model inference.
- Study analyzed 26 LLMs from seven families across subjective and objective tasks.
- Larger models show a gradual decrease in contextual entrainment effects.
- 2-4% of attention heads control contextual entrainment and can be disabled effectively.
- Disabling attention heads does not harm overall model performance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Contextual entrainment, which is a newly discovered phenomenon in large language models (LLMs), refers to the tendency of a model to assign higher probabilities to tokens that appear in its context. In this work, we extend this phenomenon from the token level to the sentence level by examining the per-token mean log-probability of a sentence instead of the probabilities of individual tokens. We investigate sentence-level contextual entrainment across 26 LLMs from seven families and two datasets, which cover both subjective and objective tasks. We find that sentence-level contextual entrainment exists. This means that the sentences in the prompt (even if they are counterfactual statements) can significantly increase their probability during model inference time. As the model size increases, contextual entrainment gradually decreases. We also find that contextual entrainment is controlled by 2% to 4% of the attention heads. Turning off these attention heads can effectively mitigate contextual entrainment without hurting the model's performance.
| Comments: | 16 pages, 3 figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.24077 [cs.CL] |
| (or arXiv:2606.24077v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24077 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yang Liu [view email]
[v1]
Tue, 23 Jun 2026 02:42:33 UTC (219 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.