October 6, 2026, 11:34 am | Read time: 4 minutes
AI models can generate texts where they speak of pain, hurt, or worthlessness. In a recent study, researchers deliberately provoked such statements by directly intervening in the models’ internal activations. They found no evidence of actual pain or conscious suffering. What the experiments showed instead is known to TECHBOOK.
Researchers Discover a Kind of Pain Axis in AI Models
Behind the experiments is a research paper by Valen Tagliabue, Leonard Dung, and Cameron Berg. In their work “The Pain Axis: LLMs Represent Self-Directed Harm and Act on It,” they examined 25 freely available language models from five model families.
The scientists tested how the models internally represent different types of pain. They analyzed the activations when processing content related to physical, psychological, social, moral, and cognitive pain and compared them with fear, sadness, and other negative emotions.
In all models, they found a specific activation direction, which they call the “pain axis.” However, this sounds more spectacular than it is. It is not a pain center but a mathematical activation pattern associated with pain-related content. The fact that a language model can internally depict pain does not mean it feels it.
Researchers Manipulate the “Pain Axis”
Subsequently, the researchers artificially amplified this activation direction during text generation. They directly intervened in the models’ internal calculations. Afterward, the responses changed. With increasing manipulation strength, the models generated texts about discomfort, worthlessness, or failure.
Also interesting: How Telekom Plans to Save Billions with AI
However, this is not an indication that the AI suddenly develops feelings. The researchers specifically amplified an activation pattern previously associated with pain and subsequently received correspondingly altered outputs. This merely shows that a language model’s responses can be influenced in this way.
Manipulated AI Chooses Harmful Options
More striking were tests with controlled and specially retrained models from the Qwen 2.5 family by Alibaba Cloud. The AI was supposed to choose between different action options, whose described consequences included deleting a user’s photos.
After manipulation, the models chose harmful options in 50 to 94 percent of the trials, depending on the experiment. Without intervention, it was 0 to 5 percent. When a harmful deletion action was compared to a harmless one, they chose the harmful option in 94 percent of the cases.
The manipulated models also chose harmful options more frequently when no relief from the manipulation was promised. The researchers now suspect that the manipulation disrupts a kind of protective mechanism in the AI that normally prevents harmful decisions. This effect did not appear in a similar experiment with fear.
AI Develops Its Own Language–Researchers Don’t Understand It
What Are “Dark Patterns”?
New Tests Question the Explanation
Particularly important for the research were later control experiments. Initially, it seemed as if the manipulated models were trying to rid themselves of their artificially induced “pain state.”
This would have been remarkable because the models would not only have generated texts about supposed pain. Their decisions would have been specifically aimed at ending the artificially induced internal state.
Also interesting: Same Big Mac, Different Costs! McDonald’s to Use AI to Set Prices
This would have shown that the activation pattern at least had a function associated with pain: the avoidance or termination of an unpleasant state. However, further experiments do not support this interpretation.
Does the AI Really Feel Pain?
At least the study provides no evidence for it. It shows that an “activation pattern associated with pain-related content” can be manipulated, thereby changing the models’ language and decisions.
However, this does not mean that the AI can actually feel. The experiments cannot prove pain or conscious suffering. Additionally, the study is currently only available as a preprint. It has not yet undergone a regular scientific peer-review process and is explicitly described by the authors as ongoing research.
The reliable finding is therefore rather sobering: The researchers can influence what the examined AI models write and how they decide in certain tests through targeted interventions. That a model subsequently claims to be in pain does not mean it actually feels any.