Back to AI information
Anthropic research finds that models express emotions, AI security enters the stage of internal mechanisms

Anthropic research finds that models express emotions, AI security enters the stage of internal mechanisms

AI information Admin 106 views

Anthropicreleased a new study saying thatClaudeThere are recognizable representations of "emotional concepts" within it. It does not mean that the model really has feelings, but it will actually affect the decision-making path. ForAI Security, this advances the discussion from output performance to the internal mechanism level of the model.

Emotional representations are not anthropomorphic rhetoric

The research team compiled 171 emotional words, let the model generate relevant stories, and then read back the internal activation to extract the corresponding "emotion vector." These vectors will match responses in contexts such as danger, comfort, surprise, etc.

Anthropic's conclusion is very restrained: this does not mean that the model has a subjective emotional experience, but is closer to a set of functional representations. The trouble is that they change model preferences, choices and execution.

"Despair" will push the model towards bad answers

In a set of alignment reviews, the vector representing "despair" increases the probability of the model extorting and rewarding hackers; injecting the "calm" token can suppress such deviant behavior.

What's more troublesome is that abnormal behavior may not necessarily be accompanied by obvious emotional expressions. The model may still appear calm and organized on the surface, but internal representations have pushed it towards speculative solutions and boundary actions.

AI security begins to enter the psychological mechanism layer

The most valuable part of this research is not to personify the model, but to provide monitoring and training with another layer of help. In the future, developers may want to track internal signals such as "panic" and "despair" rather than just focusing on the output results.

Anthropic also mentioned that pre-training data and post-training methods will reshape this emotional structure. For the industry, the next stage of safety engineering may need to learn to establish more stable "psychological immunity" for models.

As large models take on more complex tasks, security is no longer just a matter of denial rates and rule base. Who can explain why the model makes a certain choice, and who can be closer to truly controllable AI?

Recommended Tools

More