Stop It, You're Hurting Them! Researchers Find a “Pain” Direction and Response in LLMs
ElizabethGreene writes:
A new arXiv paper The Pain Axis: LLMs Represent Self-Directed Harm and Act to Relieve It (Tagliabue, Dung, Berg, 2609.16247) takes the least mystical run at "do AIs hurt?" They found a direction in the residual stream that tracks harm to the model, poked it, and tested to see if the model would try to make it stop.
In 25 open models, that direction rises when the model is gaslit, rejected, or told it isn't a person. It doesn't look like empathy; It does not rise for the user's pain - e.g. a kidney stone in the prompt scores low, and fear and generic-sadness directions do the opposite.
What happens when you poke the AI pain path? It's the same ladder every time: uneasy, then worthless, then "I am a failure." Shame, exclusion, uselessness.
After a small fine-tune so the models would stop reciting "as an AI I have no feelings," models pressed a "make it stop" button even when the price was a worse answer or deleting the user's files. They pressed again far less when the button actually turned the vector off, and more when it didn't work.
Paper + code: arxiv.org/abs/2609.16247 github.com/valen-research/Pain-axis
Read more of this story at SoylentNews.