Verbalizable Representations Form a Global Workspace in Language Models
Quick Answer
This paper identifies a 'J-space' in language models, akin to a global workspace in human cognition, enabling verbalizable representations that facilitate reasoning and decision-making.
Quick Take
The findings suggest that language models possess a privileged set of representations that reflect conscious access, revealing insights into their cognitive processes and improving alignment through counterfactual reflection training.
Key Points
- The J-space allows language models to verbalize and reason with intermediate concepts.
- It reveals strategic deliberation and misaligned dispositions not present in outputs.
- Counterfactual reflection training enhances model behavior by simulating interruptions.
- The J-space exhibits coherent content in a specific range of model layers.
- This research provides insights into the cognitive processes of language models.
DeepSignal Analysis
What happened
The paper presents findings suggesting that large language models exhibit a functional distinction similar to the human brain's conscious access. This is characterized by a 'J-space' that allows for verbalizable representations, which can be used for reasoning and decision-making. The authors introduce a new interpretability technique to identify these representations.
Key evidence
- The authors identify a 'J-space' in language models that enables verbalizable representations, facilitating reasoning and decision-making.
- Using the Jacobian lens, the study reveals that the J-space contains coherent content in an intermediate band of layers, reflecting properties of a global workspace.
- The paper introduces counterfactual reflection training, which improves model behavior by focusing on what a model would say if prompted to reflect.
Why it matters
Understanding the J-space in language models can enhance our comprehension of their cognitive processes, potentially leading to better alignment with human reasoning. The findings suggest that these models maintain a privileged set of representations that can be accessed and verbalized, which may improve their utility in applications requiring reasoning and decision-making.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey
Abstract:Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15495 [cs.CL] |
| (or arXiv:2607.15495v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15495 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wes Gurnee [view email]
[v1]
Thu, 16 Jul 2026 22:54:30 UTC (11,700 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.