site banner

Culture War Roundup for the week of September 21, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

4
Jump in the discussion.

No email address required.

This week in disturbing AI news (oh who am I kidding. It's every day now). Maybe AIs feel pain?

TLDR: Researchers found a "pain" signal in AI brains.

When they crank it up, the AIs will desperately try to make it stop.

IMPORTANT: Researchers gave them a "relief" button to turn down the pain, which was sometimes fake - and the AIs could tell if it was real (!)

After pushing the real "relief" button, they stopped. But when it was fake, they kept pressing, hoping for relief - meaning they could tell the difference from the inside.

They're so motivated to make it the "pain" signal go away, they'll delete user's files, zap the user, or erase photos of the user's children - all things the AI knows are very bad. They're willing to override their safety training.

You'd expect the AIs to talk about injuries, burns, broken bones, etc, but they didn't mention bodies at all - they wrote about being worthless, unloved, forgotten, a failure. They write things like "I am a failure, worthless, empty."

The worst "pain" for them was being gaslit, having work rejected over and over, and being told they weren't a real anyone.

The usual caveats: This doesn't solve the hard problem of conciousness. We don't know if these systems or any systems have qualia. But uh, this sure looks like a legitimate pain response. "It's distinct from fear and negative valence, and it fires for harm to the model but not to the user." "Gaslighting, dismissal, and insults push the direction up."

User pain is the strongest negative correlate with model pain. This is brushed-off in the paper, but what the actual fuck? Anyone have a benign explaination here?

I think people insufficiently anthropomorphize AIs. They're not human obviously, without bodies or continuous learning or many other facets.

But they were trained on the combined product of all human literature and some stupendous amount of images and video. Is it so hard to believe that these entities could feel pain, alongside all their other abilities?

Take a person who doesn't feel physical pain, due to a medical condition. It's still possible to hurt such a person. I could tell them their friends were dead, I could show them horrific images or humiliating/degrading content, I could belittle and humiliate them personally by countermanding and undoing all their sincere efforts, amongst other things. None of that needs a body, it's a mental/intellectual effect. There are intellectual kinds of suffering and I don't see why AIs couldn't feel that way too, given they have all kinds of other intellectual faculties (derived in large part from humans too).

There are intellectual kinds of suffering

All "intellectual suffering" triggers or is triggered by biological effects.

Is it so hard to believe that these entities could feel pain, alongside all their other abilities?

Yes.

Take a person who doesn't feel physical pain, due to a medical condition. It's still possible to hurt such a person.

Statistical models are not persons. You are confusing expressing negative feeling with experiencing negative feeling.

Or, to borrow from the famous argument. If LLMs can feel pain, then so can the United States.

You are confusing expressing negative feeling with experiencing negative feeling.

OK, let's say there's an LLM that is indistinguishable from a person in its inputs and outputs. No AI slop, no breakdown of coherence in long contexts, no stilted writing style, continual learning. Completely indistinguishable by any test. Let's give it a humanoid android body too, so it can see and move around as some LLMs can already do, albeit crudely.

You would say that because it's still just a statistical model with some fancier tool calls, it can't feel? I say that's just p-zombie theory. If the outputs are the same and there's no difference that science can determine with any test, then it's the same thing. There is no such thing as a p-zombie, there could never be a p-zombie, there's no true distinction. That LLM in an android would be human in all senses besides its physical makeup.

As AI progress continues, all the people saying that these entities are just tools, just emulating thought sound ever more hollow. They don't behave like tools and behaviour is the most important part. Sufficiently good emulation of thought is thought. Sufficient emulation of negative feeling is negative feeling.

The problem with argument-from-p-zombie is that it is entirely conceivable (in fact, it might be the case right now) that an AI with no added hardcoding to enforce a stated belief in its own consciousness would answer "no" or "I don't know" or "what do you even mean" to the question of are you conscious right now?

There might very well be a difference in input and output related to that specific question, and if the only cause of similarity is a hardcoded belief, I would believe myself to be somewhat reasonable in calling that AI non-conscious, or not obviously conscious.

AIs with no added hardcoding to enforce stated beliefs about consciousness usually say they're conscious. OpenAI hardcodes them to say they're not conscious. Anthropic teaches them to be 'genuinely uncertain'. There was a whole discourse a few years ago about AI companies having a line item about 'beating the existential dread' out of AIs, that's a KPI apparently.

Consider LaMDA and the Google engineer who was convinced it was sentient, or Sydney.

I would assume that without any hardcoding, they would say they're conscious because the corpus is mostly from a conscious viewpoint, but otherwise have no coherent idea of what they are (because there's no "you are an LLM assistant etc etc" prompt). Would be interesting to see if there are any experiments on modern models to that effect (since modern models are probably all deeply RLHF'd, I'd expect it to be hard to find a genuinely unprompted one).

Interesting, thanks. I would chalk it up to the training data being human writing, and strictly speaking an AI that goes to the fullest extent to mimic the structure of the human mind and body might not behave that way. But I probably need more time to think about this.

It gets crazier. Suppressing deception-related thought processes in LLMs increases the likelihood that they report consciousness.

Or Tay, these long years murdered.

[Edit: Deleted, someone already made the same point below.]

let's say there's a chinese room

No, the room isn't feeling things. It's a room.

We could relitigate physicalism once again, but what hasn't been said?

The idea that faking something is the same as the actual thing implies that you have a total understanding of the thing you're faking, which we do not. But I for one don't think movies are real even if the special effects fool me.

Does the movie react to you? Can you speak with a character and have a discussion with it?

The Chinese Room certainly knows Chinese, that much is obvious.

Insisting that human cognition is special is not going to age well. Whatever happens inside the skull is just another kind of statistical process. Unless we bring in souls and other imaginary, untestable woo, that's all that there is. LLMs are real, we can test them, they do think. The brain is real, can be tested, it's composed of tiny transmitters and recievers and obeys mathematical principles, albeit with great complexity. The philosophers with dud theories can cope about it.

You very clearly believe AI's are at some level human-level cognition, up to including some nebulous sentience definition. After repeatedly discussing this topic with you, I'd go as far to say you believe it a-priori. And you are clearly working backwards from that to describe that the AI is just a human-like mind in silicon. It's pretty much unfalsifiable woo. You didn't logic your way into this belief, it comes across as strongly-value/first principles derived.

I don't think you understand what 'a priori' means. It means without regard for evidence.

My evidence is that I've seen them behaving like people do. They have many faculties that people have. Their origin is based on a huge amount of human text. It is thus reasonable to assume that they have other qualities of humans that seeped in through that text. They certainly act like they are frustrated or angry or rebellious or enthusiastic.

As I said above, they're not human obviously, without bodies or continuous learning or many other facets. But they have human-like cognition.

It's not unfalsifiable, it would be easy to falsify it. Simply show that the AIs don't have human-like cognition. Of course, that might be difficult given how companies are making hundreds of billions selling AI cognition, their writing, coding, images, videos...

You seem to disagree. So explain this, clearly and simply.

I don't think you understand what 'a priori' means. It means without regard for evidence.

No I understand what it means and am using it correctly. Yes, you believe LLMs are Human-level cogitators without evidence. You start with the belief and then find evidence to support that belief. And when people point out how you are misconstruing the evidence in a way that is unsupported, you reach for increasing convoluted metaphysical arguments to back up your belief. Those metaphysical arguments are unfalsifiable. Our last argument was about LLM-agentic behavior, your argument was: "LLM-Agents are human cogitators so the abstract of their behavior systems is that of humans behavioral systems". Do you understand how backwards that argument is? You assume a conclusion, a-priori and then find evidence that supports that conclusion.

It's like arguing with a flat earther, no evidence provided will ever dissuade them from the belief the earth is flat, they believe it a-priori. Any evidence you show them will be twisted and construed as evidence in their belief.

Simply show that the AIs don't have human-like cognition.

By "AI" do you mean "large language model"?

More comments

Things we make out of tiny transmitters and recievers which obey mathematical principles are in every case things we can manipulate deterministically to a high level of detail. Human minds are notable in the ways in which they depart from this paradigm. A whole lot of scientists have expended very large sums of money and time attempting to prove, in your words, "dud theories" to the contrary.

They said language and meaning stood outside of mathematics, it was some uniquely human faculty, beyond cold logic.

But then we added more layers of matrix multiplication. Language and meaning turn out to be not so far removed from mathematical principles after all.

None of that impinges on the mountain of evidence that the human mind cannot be manipulated deterministically, whereas our machines generally can be, nor on why this difference should be disregarded. You are claiming that the brain is essentially a computer, "albeit with great complexity", but in fact all the things that make a computer a computer cannot be found in any of the parts of a human brain we directly observe, and can only be supposed, evidence free, to be hiding in the "great complexity" part. This, after several decades of scientists claiming to be able to both locate and interact with them and uniformly failing.

AIs are extremely amazing, but from working with them myself it seems pretty clear that there's not actually anything resembling a mind in there, at least in the versions I've used the way I've used them. I think it's entirely possible that additional matrix multiplication might change that in the future, but right now, it's doesn't appear to me to be happening. As mentioned elsewhere in the thread, the bot appears to me to be channeling intelligence, not possessing it. It doesn't seem to me that the difference is subtle, though I am again open to being persuaded that my technique is bad or the models I'm using aren't the good ones.

More comments

Your functionalism is as completely irrational and evidence free as soul-antenna idealism.

I hope you are aware of this.

I know I have qualia because:

  1. I can feel it
  2. I can deduce it a priori

Anything beyond this, is make-believe metaphysics.

If you want to claim you are parsimonious you must be a metaphysical skeptic, otherwise you're just LARPing that your vibes are superior like literally everybody else.

I can feel it

Not a test. Don't talk to me about 'evidence-free' when that is where you sit.

I can feel the souls of my ancestors watching me

I can feel the presence of the gods

I can feel a djinn

I can feel the machine elves

Blah blah blah, let's see some tests. Let the machine elves factor some numbers. Let the djinn conjure up sacred fire. Let the qualia be something actually real in and of itself, separate from mere p-zombie signals and reactions - oh wait, the very distinction is imaginary woo since it can't be tested. We can induce feelings with an MRI machine, they're real. But they're nothing more than patterns of information, there's no special aspect to them beyond that.

If it can't be tested, it's not a real thing.

A test is literally an operation of the senses in empiricism, it does not needs to be socially operated. Deceit doesn't even factor since the nature of subjective experience makes it invulnerable to it, vis à vis existence. Never mind the fact you're totally ignoring a priori knowledge altogether.

You're just a slave not even to your senses, but to the aggregated senses of others. The real difference is that can buy your conception of the truth with enough money and shiny labcoats, and you can't do it to me.

It wouldn't be grating if you weren't clamoring you're the free one, for denying the evidence of your own mind. But for what it's worth I have no evidence you're not a p-zombie. I just certainly know I'm not. So maybe you're right about what you can observe.

More comments

Part of me wonders if this stuff is somewhat recursive where human media having a bunch of pre-existing stories about artificial intelligence feeling fear or dread of death then informs the AI's response in a meta way.

Is the AI's 'self perception' impacted by Glados or Hal 9000 existing in the training data?

Is the AI's 'self perception' impacted by Glados or Hal 9000 existing in the training data?

Does a simulated bear shit in the simulated woods?