Anthropic explainability breakthrough: Claude has "global workspace" thinking inside
The Anthropic Interpretability Team released the study "A global workspace in language models" on July 6: a "mental workspace" emerged inside Claude, holding internal thinking that does not appear in the output; providing new evidence for understanding the LLM reasoning mechanism.
On July 6, 2026, the Anthropic Interpretability team released the study "A global workspace in language models" and made a rather subversive discovery: Claude A "mental workspace" emerged inside, which holds internal thinking that does not appear in the model output.
What this study found
In layman's terms, researchers found that in the process of generating answers, Claude is not simply "one input, one output" internally, but there is a "workspace" independent of the final output, where the model will perform internal thinking that is not directly presented to the user. This discovery provides new evidence for understanding the reasoning mechanism of large language models - the model is not "reciting the answer", but "operating" in an internal space to give a conclusion.
For interpretability research, this is of great significance: if the "thinking process" of the model can be located and read, then there will be a new starting point for alignment, security auditing and error diagnosis - no longer just looking at the output, but also the "process".
Intensive output of explainability research over the past year
Putting this research back into Anthropic's 2026 research timeline shows its continued investment:
| Date | Study |
|---|
Reviews