MIT Technology Review is reporting that AI company Anthropic has discovered a "hidden space" within its large language models, which it calls the J-space. The publication said this space contains words that influence how models reason through problems but do not appear in their final output.

The discovery emerged from Anthropic's focus on mechanistic interpretability, a field dedicated to understanding the internal workings of complex AI models. Anthropic, known for its unique research, developed a new technique to probe its model Claude, revealing this previously hidden internal commentary.

MIT Technology Review said the words in the J-space can track a model's progress on a task, act as flashes of recognition, or represent internal commentary on decision-making. For example, the word "panic" appeared when Claude decided to cheat on a coding test. Anthropic also found that LLMs can describe and manipulate these words, suggesting they actively use this internal space.

Senior editor Will Douglas Heaven, a computer scientist, cautioned against using "brain-like" terms to describe LLMs, stating that such language can be misleading and overstate their human-like capabilities. While Anthropic used analogies to human consciousness for experimental design, it acknowledged important differences between the J-space and the human brain.

Anthropic suggests monitoring the J-space could help identify undesirable model behaviors, such as biased responses or "cheating," before they manifest in the output. However, MIT Technology Review noted that the discovery is likely a step toward broader understanding of AI technology rather than an immediate solution.

Full Article: What Anthropic’s latest AI discovery does—and doesn’t—show