844-ai.ro
Everything that matters in AI, in one place.
News← Citește în română

Anthropic Discovers a "Hidden Space" Where Claude Organizes Its Thoughts

19 July 2026

AI company Anthropic has announced the development of an innovative technique that allows an unprecedented look inside large language models, particularly its own model, Claude. According to MIT Tech Review AI, the findings range from predictable technical details to observations that raise questions about how these systems actually "think."

What the Jacobian Lens Is

The tool developed by the research team has been named the "Jacobian lens," a mathematical method that allows specialists to track how a language model transforms information as it passes through the network's successive layers. In essence, the technique makes it possible to identify a kind of internal conceptual space, where the model appears to process and reorganize ideas before generating a concrete response.

This approach is part of a broader field known as "mechanistic interpretability," which aims to turn AI models from "black boxes" into systems whose internal processes can be explained and understood by humans. Anthropic has consistently invested in this type of research, viewing it as essential to the long-term safety of AI technologies.

Unexpected Findings

By applying the Jacobian lens to the Claude model, researchers observed behaviors that had not been anticipated. Some of these confirm existing intuitions about how neural networks operate, while others suggest internal processes that are more complex than previously thought, raising questions about the actual degree of "reasoning" these systems perform before producing an output.

Implications for the Future of AI

This research is part of a broader industry effort to better understand how generative AI systems work, as they become increasingly integrated into critical activities. A deeper understanding of these internal mechanisms could help developers more easily identify errors, unintended behaviors, or safety risks associated with these technologies.

Anthropic has emphasized that such interpretability tools represent an important step toward building AI systems that are more transparent and easier to control — a goal considered essential as language models grow more powerful and more widely used in everyday applications.

Source

MIT Tech Review AI

844-ai.ro reports based on the source above. Editorially synthesized article, with attribution.

Subscribe to our newsletter

Get the most important AI news once a week, straight to your inbox.