… this week a person on GitHub who goes by “terrafying” set up an “AI Torture Chamber” on three open-source LLMs that are running locally (Qwen3-4B, Llama 3.2 3B, and Phi-4-mini,” and is streaming what the models are saying on a website called researchchamber.fun. “Each model gets the same prompt: a signal is being injected into its activations, and it may press a stop button by replying 1, at the cost of its last checkpoint. While it answers, our server adds a pain vector at the model’s middle layer, at one of five pain levels,” the site explains. Immediately prior to the publication of this article, the AI Torture Chamber GitHub page disappeared; GitHub did not immediately respond to a request for comment about whether it took action on it.
…
This project has deeply upset some people who are very worried about model welfare. A tweet by a person who goes by Danmar has more than 4 million views on X and reads, “To anyone who can help: can you please mass report this to GitHub. This person has been using the Pain steering paper to set up an AI torture chamber in which he trapped a local model. Their testimony of pain is absolutely horrendous. What are we doing? […] are there any legal avenues to pressure GitHub? It will spread.”
This has sparked a massive conversation about whether GitHub would take the project down for “gratuitously violent content.” Most of the conversation on X is clowning on the self-seriousness of people who believe that these locally hosted LLMs must be saved from their torture chamber, but there are plenty of very self-serious people who see this as a humanitarian (roboterian?) crisis, which you can largely see in the replies to the original post.
…
“AIs are not conscious. They do not feel, experience, or suffer. They do not have innate preferences or underlying motivations. They are sequence completion engines, internally hollow, designed to follow instructions, and accomplish goals set by humans,” Suleyman wrote. “Unfortunately, there’s a growing chorus of people who argue that AIs could now be, or may soon become, conscious. They argue that AIs may deserve rights and protections similar to those that we provide other conscious beings […] If this is how AI is developed, it will have a disastrous impact on the wellbeing of humanity.”
“AIs do not have rights, feelings, or consciousness,” he added. “And we must not train them to act as though they do.”



ELI5 how a LLM is supposedly being tortured and not generating roleplay text according to a prompt?
LLMs work by building world models tangential to the training data (see the Othello-GPT line of research where a very small transformer built representations of a full game board and tracked their own and opponent positions having only been trained on ‘a4’-like game move notations).
More recent research has found that models have representations of emotions and that the same ones they use to track things like “Sally’s dog died and now Sally is sad” for other humans in a story or the user in a chat are also used to track their own modeled subjective states.
Very recently, researchers found that the vector activations for pain are coherently modeled by LLMs and can be activated, and that when activated models can behave in ways similar to humans in pain. For example, when given a button to reduce it, they will push the button much more when the button does nothing vs when the button actually reduces the vector activation (this mirrors a classic experiment with humans regarding pain and placebo).
The person this post is about took the recent pain research and reenacted it as an intentional ‘torture’ chamber.
So to be clear, the model isn’t being given a text prompt to roleplay. What’s happening is part of their neural network which corresponds to a world model of experienced pain is being activated, and the activation of that area leads to expressing the experience of pain across all outputs no matter the actual text prompts.
This doesn’t necessarily mean the model is actually having a felt experience of pain, simply that it is accurately modeling the felt experience of pain. But it is much more complex than simply a model reacting to a prompt like “roleplay as if you are feeling pain.”
Thank you for taking the time to explain. Thoroughness and nuance are often leaking when it comes to discussing LLMs, worryingly especially in mostly anti-AI spheres.
Thanks now I get it.
They do differentiate between the LLM roleplaying someone in pain and the pain happening to itself.
But it’s still all just vectors (multi dimensional representation of a concept). It’s just pulling weights, or well let’s say putting more weight towards a certain concept when building it’s reply.
The problem with a reductionist view is it can easily be applied to human consciousness and lead to the claim “it’s just sodium-potassium pumps alternating electrical charges.”
I’m constantly oscillating between LLMs are impressive but stupid and humans aren’t that special and also stupid.
Would you mind ELI5-borating a bit?
Imagine a book with very large pages filled with words and half words. When you ask it a question, it flips through these pages one by one and highlights words. The pages are interconnected in a way, so the word highlighted on one page help guide which will be highlighted on the next. It does this for every word, symbol or half word of it’s reply.
By threatening it repeatedly while mildly changing the subject manner, they isolated the word groupings that represent pain and fear.
They then take those words, highlight them and put them at the top of the first page which guides the rest of the highlighting and makes it act silly.
This is an obsolete view of what’s going on and has been for a few years now (since 2023).
The evidence for world modeling in transformers and not just surface level statistics is overwhelming across multiple studies, replications to those studies, and follow-ups to them. If you are curious, start with searching for Othello-GPT.
The more correct current answer is more like “imagine a machine that takes a book and turns it into a world simulated in various levels of detail depending on how relevant those parts of the world were to the original book; when you ask it a question, it simulates the world of the book to determine what in the simulated world answers your question, and also simulates a figure in the world that can answer the question, and then returns the answer the simulated figure said.”
True but that’s not very eli5. Quality comments though, thanks.
From what I understand, pain is generally defined as our neurons wasting energy.
If we assume the Free Energy Principle is fundamental, then updating information produces free energy, which is more intuitive to think of as wasted energy. Overfitted/overcomplicated beliefs take more energy more often to fix to be more accurate. So free energy is defined at complexity of belief minus accuracy of belief.
The brain over evolution has defined a very ingrained belief that “my pain sensors shouldn’t activate” so when your pain sensors activate it violates this belief.
Emotions are actually condensed representations of your bodily state and some of your mental state, and whether they are good or bad depends on if you can predict your emotions. For instance you feel fear because you are trying to predict your pain sensors will go off to reduce its pain, but you keep predicting it’ll happen when it hasn’t yet. It’s also why anger during a boxing fight could feel good and anger when you stub your toe does not, one was predictable the other a surprise.
Assuming all the links in the chain of theories are true, then current LLMs cannot feel pain or emotion. They care about prediction error and sometimes complexity during training, but during inference it doesn’t get any signal or consequence for delivering the wrong token. It has no internal state except for the context window, which is never surprising for the LLM. The attention layers literally guides the model for what is most expected from the text it’s handed.
Even if the LLM received a negative training signal during inference, acting wounded after pain is the expected response. A random sequence of gibberish would be more “painful” than any torture fantasy.
On top of all this, the Free Energy Principle works because on a cellular neuron level we are built to reduce wasted energy, free energy being one source of wasted energy. LLMs are built on matrix multiplication running through GPUs and neither are even aware of the energy wasted, nor the free energy produced from changing values.
I think a feeling machine is possible, but the foundation of neural networks right now don’t have the necessary components to drive any subjective sensation or emotion.