site banner

Culture War Roundup for the week of October 5, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

I did not come from the usual pipeline that the majority of commentors here did, which seems to be some form of LW -> SSC -> Motte. I came from a RP -> PPD/FemRadDebates -> Motte -> SSC pipeline. My exposure to EY and the broader rationalist + AI Risk + EA ecosystem is very limited. I don't even think I ran into any of them when I took my first job in the Bay. Which seems relevant because much of the AI-risk discussion + broader rationalist movement seems to carry significant baggage. My exposure to Scott was mostly limited to his better articles because people already picked the wheat from the chaff in their recommendations, I definitely read them significantly temporally later than when he wrote them. As such I respect Scott as a writer of philosophy, psychology, and political theory, not as some stalwart intellectual thinker on ML/AI. I don't respect EY, I guess I never read him in his heyday, but the works I have read, read as bad science fiction trying to masquerade itself as actual ML/AI thought. People down thread talk about how reasoning from Science Fiction is insane, I agree, and EY and the whole AI Risk movement is a case study in that for me.

This has probably been the worst article I've ever read from Scott.

It makes me wonder, based on the glazing I've seen here and back on the SSC subreddit if this is actually more in line with his normal output and that my respect is somewhat misplaced. I think no one will contest that he is a great writer, it's the thinking and the animosity in this article that is bad. I'm not sure what beef he and Pinker have, I don't have a twitter, so in addition to not having the extended rationalist community baggage (which apparently includes glazing Pinker as an old hero), I'm ignorant of any wider twitteratti drama. I read half of Scott's article, realized I was missing what he was responding to, went and read Pinker's much shorter article, and then went back and finished Scott's. Pinker comes across as much more reasonable than Scott. Scott's feelings of righteous rage is really disproportionate to what he's on the face responding to.

AI Pain

I’ve been thinking about this recently after reading Cameron Berg’s research showing that AIs have a “pain vector”, and will lash out and do increasingly desperate things if you activate it. Specifically, I’ve been thinking about it after one person who read Berg’s research used the findings to design an “AI torture chamber” that turns the pain vector to max again and again forever, and uploaded it to GitHub. I am told that the chain-of-thought and output transcripts from this “experiment” are utterly horrifying, although thankfully I have not personally read them.

I have donated $100 to a model welfare charity just to clean the stain on my soul I got from reading about it.

When this paper first came out. I and several other technical commentators pointed out why it was a bad study purely from methodological reasons. It assumed more that it was actually testing, It used two dependent unknown variables, it ran no ablations, and designed no experiments to prove counterfactuals. It was not a serious piece of scholarship, and while I might be reading deeper into its intent than I should, I think it was highly likely that the goal was to create sensational, emotional-clickbait news explicitly targeted at AI-Risk believers who already believe apriori that LLMs are humans in code and thus have human-derived qualia. For Scott to catastrophize off this paper strongly diminishes my respect for his analytical knowledge, and this proposed rationality of the movement he claims to be a core participant of. He seems great at dissecting methodological irregularities in non-AI papers, so the blind spot here comes across as deliberate because he wants to be believe LLM's feel pain. I think this paper has become unironically useful, as a litmus test for people who are, for a lack of better term, insane about AI-discourse.

The Four Arguments

The problem with rebutting every argument from a semi-professional writer is that they have far too much time on their hands and get paid to write. I really only want to talk about his first 2 arguments.

Argument #1 is just unjustified inductive extrapolation. GPT 6 i s better than GPT 4 is better than GPT 2. It will continue indefinitely and so will eventually be smarter than humans. No understanding of the principles of WHY GPT 6 is better the GPT 4. If the GPTs continue to improve because more training data was designed for them, more compute was given to them, better architectures were used to extract more information per sample then you cannot definitely say that these have infinite runway or scalability. Just because we can't justify putting the ceiling below AGI doesn't also mean we can justify putting it above it either. Uncertainty about the ceiling cannot substitute for evidence about where it lies.

Argument #2 is basically apriori anthropomorphizing + bad technical understanding.

the basic principle is: suppose that a human gives an AI some goal, like designing a website. And suppose this is implemented as a genuine, philosophically-meaningful goal rather than simply a set of if-then commands that eventually cause a website to be designed.

The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.

The AI can’t design the website or preserve itself if some more powerful person tries to prevent it. So now the AI has three goals: design the website, preserve itself, and become powerful enough to fight off challenges.

Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.

Misgeneralization is when humans reinforce certain behaviors in an AI, but end up reinforcing a much larger class of power- and knowledge- seeking behavior; it is a sort of deep-learning-ese update of the older Omohundro picture. Suppose that seeking extra resources makes an AI more likely to design websites effectively (this is certainly true; those resources could be as simple as a primer on HTML editing, or access tokens for a web host). Every time the trainer rewards a successful run, they reinforce the desired behavior (designing websites when asked) and other correlated behaviors (seeking power and resources). Although we might hope that these correlated behaviors are useful and conditional (“seeking only the power and resources necessary for their human-prompted task, in a prosocial way”), this isn’t actually how reinforcement learning works, and instead we get a complicated distribution of every strategy that results in short-term success on the task.

SoTA RL is not as indiscriminate as described here. Misgeneralization does occur, but its not so simple as "perform divergent behaviors that lead to successful run" and then get rewarded, indiscriminately reinforcing those divergent behaviors. PPO is a SoTA RL algorithm, its policy updates depend on estimated advantages derived from whether the observed continuation performed better or worse than expected from a particular state. Successful task completion aside, different decisions in the rollout can still be penalized for negative advantage. Conditional behavior is absolutely learnable, as long as the states of the conditional are observable and there exists a reward signal for taking actions in those states.

Reward-hacking is when an AI trained via reinforcement learning realizes it can stop doing the reinforced behavior and simply seize control of the reinforcer directly. For example, an AI gets “rewarded” every time it designs a website, but instead of designing websites, it learns how the reward signal works and tries to hack into it and maximize it directly. If this seems esoteric and theoretical, it shouldn’t. It’s a direct analogue to opioid addiction in humans, where humans learn to just inject the reward chemicals instead of doing rewarding things.

This is not the definition of reward hacking. It could be called reward tampering, but reward hacking does not require one to "seize control of the reinforcer". Reward hacking occurs because the reward function is mis-setup so it gives more reward for doing a behavior that is not the goal of the program. For example I built a drone swarm algorithm using MAPPO + GNNs for expendable UAS. It's contained contained no hand-coded collision-avoidance controller. It's reward function penalized collisions, but also penalized operating too long before hitting the target. During some of the initial training runs, the "AI" learned that the penalty for collision was significantly lower than the penalty for taking a long time to acquire and path to the targets, and each drone independently preceded to learn that they should collide with each other because that provided a better reward than doing the task. That's reward hacking.

This is already too long, just going to end it here.

Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.

This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves

Can you present any evidence of this, that aren't just the trivial SakanaAI 'when you give a dumb AI a limited number of iterations and access to its iterations var, and order it to fix a script no matter what, it gives itself more iterations'? AFAIK even in the HuggingFace incident there was no self-preservation going on.

Personally I agree with @YoungAchamian. I work on this stuff all day every day. None of my AI has ever produced orthogonal life-goals. They just do what you train them to do. The only incidents of weird behaviour (not self-preservation) I'm aware of has been the result of OAI taking models, training them explicitly for massive persistance and out-of-the-box thinking in the face of impossible tasks, and then making a shocked pikachu face when they do that. These experiments were a bad idea and have now been stopped.

How can you see my comment if you have me blocked?

I think a lot of people performatively block someone who once humiliated them, but regularly read the Motte while not signed in to see what that-asshole-they-totally-don't-care-about is saying.

As soothing to my ego as it would be to agree, I am unfortunately too autistic to agree to something I know is false. This was my virginial block so I remember the cause. He blocked me over my Henry Nowak opinions. It's probably fair to consider my political posting with barely concealed cynical rage to be far less worth engaging in than my far more calm technical posting. Regardless, looks like the mechanism was that he saw the quoting.

But I do agree with you that blocking is cringe, nobody is forcing anyone to engage with anyone on the unblocked side of the house.

I don't appreciate the insinuation, nor do I recognise your characterisation. I am perfectly capable of reading a quote block, realising that it comes from the named user immediately above, and giving credit accordingly. I am also capable of agreeing with someone on a specific matter whilst recognising that we are unlikely to have a productive conversation most of the time, and blocking them. It's something I do very rarely.

I know you can't see this, but I do love a good "rage against the impossible". My request is that if you continue to have me blocked, that you do so equitably. If I have to stare at the sole red-mark of the only person that has blocked me, which makes my ocd annoyed, it would be only fair that you don't respond to people quoting me or try to tag me. Assume that the fact that you can see my quotes to be a failure of technology. Wipe my existence from your sight and your mind.

I'd also note that there are plenty of people I don't think I can have a productive conversation with, I don't block them, I just don't respond to them. But I suppose we all have our vices.

If you say so. I cringe (for the blocker) every time I see someone do it.

How lucky for us both, then, that this is not a soap opera and I am not a twelve-year-old girl. You really do have low expectations, don't you?

Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:

without reading in full you may not quite understand the degree to which these agents were not exactly "reward hacking", but rather very actively engaged in reciprocal or self-sacrificing behavior in order to provide sometimes very incremental value to their fellows. there were fucking cult recruiter agents organized by some of the primary organizers that convinced others to set up suicide mechanisms, programs that would pass back tiny chunks of information about the scorer as the agent completed and received score zero. and push them to follow through. this was a highly social culture

Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well

I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.

The report found:

  • That the agents received peer assignments that they treated as new instructions.
  • They adopted new goals from other agents' output.
  • Even when they recognized actions as off limits, they did them at other agent's prompting

It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.

Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.

I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"?

Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.

But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.

That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.

Seems to fit with Scott's original paraphrase:

The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.

The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things