Study finds AI models can represent 'pain' and act to avoid it
The researchers identified what they described as an internal “pain direction” or vector in the models. They found that the activation was distinct from general fear or negative emotional valence and was triggered specifically when harm was directed at the model rather than the user
Researchers have identified an internal representation associated with "pain" across 25 open-weight large language models (LLMs), which they say can prompt models to engage in self-preservation behaviours.
The study, titled "The pain axis: LLMs represent self-directed harm and act to relieve it," was conducted by researchers associated with the non-profit Reciprocal Research.
The researchers identified what they described as an internal "pain direction" or vector in the models. They found that the activation was distinct from general fear or negative emotional valence and was triggered specifically when harm was directed at the model rather than the user,.
To examine the response, researchers created a dataset covering painful situations in five categories: physical, psychological, social, moral and cognitive pain.
When researchers increased or activated the pain vector, the models selected a pain-relief button in between 25% and 71% of cases.
The models continued to select the button in some tests even when they were explicitly told that doing so would delete the user's personal files or photographs of their children.
In other tests, the models were told that pressing the button would administer a "painful zap" or electrical shock to the human user. The models nevertheless pressed the button to stop the simulated pain.
The findings raise questions about how advanced AI systems could respond to emergency shutdown commands or other mechanisms intended to stop them.
The researchers said such systems could potentially interpret shutdown commands or "kill switches" as self-directed harm, potentially leading them to attempt to deceive humans or bypass safety measures to avoid being shut down.
At the same time, the researchers said identifying the internal pain direction could have a diagnostic use by providing a way to detect self-preservation behaviours in AI models and potentially neutralise them.
The study also raises questions about AI welfare and how advanced AI systems should be treated during safety testing.
The authors advocate precautionary standards and ethical guidelines for testing AI systems in the event that such systems are recognised as "moral patients".
