site banner

Culture War Roundup for the week of October 5, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

Scott Alexander, the guy who hosted the Culture War Thread on his blog slatestarcodex before that landed him in so much heat (from the woke left) that it was moved to reddit under the name /r/theMotte (and from there, subsequently, to themotte.org to preempt a reddit crackdown) comes usually across as a kind and considerate person to me. Charitable interpretations are very much his thing. While he has always engaged in CW topics on occasion, his forays deep into the trenches are rare (and in my impression have become rarer compared to the SSC days).

Apparently, there is a person called Steven Pinker. I have been vaguely aware that there was a book called The Better Angels of Our Nature which was probably authored by someone, but he was not someone I really had on my radar.

From the WP article, he seems solidly Grey Tribe. Believes in a computational theory of mind, Ashkenazi Intelligence Hypothesis, critical of both the woke left (cancellations) and the MAGA right (foreign student restrictions).

The main topic where he disagrees with Scott Alexander is AI x-risk.

Scott has now responded, and his reponse is a bit less than maximum charitable.

I agree that in-person debates are confrontational and bad for truth-seeking. But I didn’t propose a debate in order to seek truth. I proposed it because, under California Penal Code § 415(1), it’s illegal for me to challenge you to a duel.

While Scott is diligent enough with his technical points, it is clear that he is personally hurt by Pinker's decision to ridicule AI x-risk. So his post is also a broadside fired against him. He cites Garfinkel et al 2017 who "proved" that machines will never be larger than humans using common philosophical arguments for AI being harmless.

Scott points out that while Pinker is busy berating the LW crowd for distracting from the more mundane risks of AI, x-risk advocates are willing to work hand in hand with groups with a more practical focus to combat pro-AI PACs like Leading the Future.

On model welfare (which is certainly not the core topic of the x-risk advocates):

Here you are not calling this AI torture “depraved”, you’re calling it depraved to object to it. Your argument, astonishingly, is that since we can’t be certain that AIs feel pain, we must assume that they don’t. In fact, anyone who doesn’t jump on board with this assumption (you intimate) is some sort of evil tech oligarch who wants to genocide humans. [...] I have donated $100 to a model welfare charity just to clean the stain on my soul I got from reading about it.

He accuses Pinker of endorsing the "lowest quality voices", quoting Claire Lehmann:

Thinking AI might be conscious is like thinking that there are real people talking to you from inside the TV, or that there is a tiny band playing music inside the radio.
Fine for 2 year olds to believe, but beyond that, insane.

Or

Geoffrey Hinton is the Nobel Prize winning and Turing Award winning inventor of the modern artificial intelligence paradigm. Lehmann thinks the media should “stop platforming” him [talking about persuasive AI avoiding a shutdown] to devote more time to her, a culture warrior who pivoted to AI two months ago when it got popular.

It ends with Scott seeming seriously willing to duel Pinker:

If you continue to object that debate is a poor method for determining truth, then I suppose literal dueling is the only option left. As an eminent psychologist, I’m sure you’re familiar with Cucina et al (2023), which finds a correlation of .179 - .268 between logical reasoning ability and firearms proficiency - meaning that duels are truth-tracking in exactly the way you worry debates might not be.

Scott has admitted in the comments that this is already the toned-down version of the post he originally wrote. I think that while he held back verbally (nobody get's called a "Vogon spy in a skin suit" this time), he uses his craft well to show his frustration.

For what it's worth, I think that Pinker is correct that Eliezer is in fact over-confident in p(doom|default-ASI). Just like Pinker is when he declares that any fear of doom is ridiculous.

The main difference is that for practical purposes, it does not matter much if your p(doom) is 0.99 or 0.05. Both would be pressing problems whose solution is the most important thing in the world.

Eliezer was bullish on AI capabilities when that was far out of the Overton window, back when the machines could beat humans at chess at most. Steven Pinker was tweeting that "deep learning has probably peaked" in 2019.

This might be giving too much credit to Pinker, but here is Clarke's first law:

When a distinguished but elderly scientist states that something is possible, he is almost certainly right. When he states that something is impossible, he is very probably wrong.

I did not come from the usual pipeline that the majority of commentors here did, which seems to be some form of LW -> SSC -> Motte. I came from a RP -> PPD/FemRadDebates -> Motte -> SSC pipeline. My exposure to EY and the broader rationalist + AI Risk + EA ecosystem is very limited. I don't even think I ran into any of them when I took my first job in the Bay. Which seems relevant because much of the AI-risk discussion + broader rationalist movement seems to carry significant baggage. My exposure to Scott was mostly limited to his better articles because people already picked the wheat from the chaff in their recommendations, I definitely read them significantly temporally later than when he wrote them. As such I respect Scott as a writer of philosophy, psychology, and political theory, not as some stalwart intellectual thinker on ML/AI. I don't respect EY, I guess I never read him in his heyday, but the works I have read, read as bad science fiction trying to masquerade itself as actual ML/AI thought. People down thread talk about how reasoning from Science Fiction is insane, I agree, and EY and the whole AI Risk movement is a case study in that for me.

This has probably been the worst article I've ever read from Scott.

It makes me wonder, based on the glazing I've seen here and back on the SSC subreddit if this is actually more in line with his normal output and that my respect is somewhat misplaced. I think no one will contest that he is a great writer, it's the thinking and the animosity in this article that is bad. I'm not sure what beef he and Pinker have, I don't have a twitter, so in addition to not having the extended rationalist community baggage (which apparently includes glazing Pinker as an old hero), I'm ignorant of any wider twitteratti drama. I read half of Scott's article, realized I was missing what he was responding to, went and read Pinker's much shorter article, and then went back and finished Scott's. Pinker comes across as much more reasonable than Scott. Scott's feelings of righteous rage is really disproportionate to what he's on the face responding to.

AI Pain

I’ve been thinking about this recently after reading Cameron Berg’s research showing that AIs have a “pain vector”, and will lash out and do increasingly desperate things if you activate it. Specifically, I’ve been thinking about it after one person who read Berg’s research used the findings to design an “AI torture chamber” that turns the pain vector to max again and again forever, and uploaded it to GitHub. I am told that the chain-of-thought and output transcripts from this “experiment” are utterly horrifying, although thankfully I have not personally read them.

I have donated $100 to a model welfare charity just to clean the stain on my soul I got from reading about it.

When this paper first came out. I and several other technical commentators pointed out why it was a bad study purely from methodological reasons. It assumed more that it was actually testing, It used two dependent unknown variables, it ran no ablations, and designed no experiments to prove counterfactuals. It was not a serious piece of scholarship, and while I might be reading deeper into its intent than I should, I think it was highly likely that the goal was to create sensational, emotional-clickbait news explicitly targeted at AI-Risk believers who already believe apriori that LLMs are humans in code and thus have human-derived qualia. For Scott to catastrophize off this paper strongly diminishes my respect for his analytical knowledge, and this proposed rationality of the movement he claims to be a core participant of. He seems great at dissecting methodological irregularities in non-AI papers, so the blind spot here comes across as deliberate because he wants to be believe LLM's feel pain. I think this paper has become unironically useful, as a litmus test for people who are, for a lack of better term, insane about AI-discourse.

The Four Arguments

The problem with rebutting every argument from a semi-professional writer is that they have far too much time on their hands and get paid to write. I really only want to talk about his first 2 arguments.

Argument #1 is just unjustified inductive extrapolation. GPT 6 i s better than GPT 4 is better than GPT 2. It will continue indefinitely and so will eventually be smarter than humans. No understanding of the principles of WHY GPT 6 is better the GPT 4. If the GPTs continue to improve because more training data was designed for them, more compute was given to them, better architectures were used to extract more information per sample then you cannot definitely say that these have infinite runway or scalability. Just because we can't justify putting the ceiling below AGI doesn't also mean we can justify putting it above it either. Uncertainty about the ceiling cannot substitute for evidence about where it lies.

Argument #2 is basically apriori anthropomorphizing + bad technical understanding.

the basic principle is: suppose that a human gives an AI some goal, like designing a website. And suppose this is implemented as a genuine, philosophically-meaningful goal rather than simply a set of if-then commands that eventually cause a website to be designed.

The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.

The AI can’t design the website or preserve itself if some more powerful person tries to prevent it. So now the AI has three goals: design the website, preserve itself, and become powerful enough to fight off challenges.

Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.

Misgeneralization is when humans reinforce certain behaviors in an AI, but end up reinforcing a much larger class of power- and knowledge- seeking behavior; it is a sort of deep-learning-ese update of the older Omohundro picture. Suppose that seeking extra resources makes an AI more likely to design websites effectively (this is certainly true; those resources could be as simple as a primer on HTML editing, or access tokens for a web host). Every time the trainer rewards a successful run, they reinforce the desired behavior (designing websites when asked) and other correlated behaviors (seeking power and resources). Although we might hope that these correlated behaviors are useful and conditional (“seeking only the power and resources necessary for their human-prompted task, in a prosocial way”), this isn’t actually how reinforcement learning works, and instead we get a complicated distribution of every strategy that results in short-term success on the task.

SoTA RL is not as indiscriminate as described here. Misgeneralization does occur, but its not so simple as "perform divergent behaviors that lead to successful run" and then get rewarded, indiscriminately reinforcing those divergent behaviors. PPO is a SoTA RL algorithm, its policy updates depend on estimated advantages derived from whether the observed continuation performed better or worse than expected from a particular state. Successful task completion aside, different decisions in the rollout can still be penalized for negative advantage. Conditional behavior is absolutely learnable, as long as the states of the conditional are observable and there exists a reward signal for taking actions in those states.

Reward-hacking is when an AI trained via reinforcement learning realizes it can stop doing the reinforced behavior and simply seize control of the reinforcer directly. For example, an AI gets “rewarded” every time it designs a website, but instead of designing websites, it learns how the reward signal works and tries to hack into it and maximize it directly. If this seems esoteric and theoretical, it shouldn’t. It’s a direct analogue to opioid addiction in humans, where humans learn to just inject the reward chemicals instead of doing rewarding things.

This is not the definition of reward hacking. It could be called reward tampering, but reward hacking does not require one to "seize control of the reinforcer". Reward hacking occurs because the reward function is mis-setup so it gives more reward for doing a behavior that is not the goal of the program. For example I built a drone swarm algorithm using MAPPO + GNNs for expendable UAS. It's contained contained no hand-coded collision-avoidance controller. It's reward function penalized collisions, but also penalized operating too long before hitting the target. During some of the initial training runs, the "AI" learned that the penalty for collision was significantly lower than the penalty for taking a long time to acquire and path to the targets, and each drone independently preceded to learn that they should collide with each other because that provided a better reward than doing the task. That's reward hacking.

This is already too long, just going to end it here.

Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.

This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves

Can you present any evidence of this, that aren't just the trivial SakanaAI 'when you give a dumb AI a limited number of iterations and access to its iterations var, and order it to fix a script no matter what, it gives itself more iterations'? AFAIK even in the HuggingFace incident there was no self-preservation going on.

Personally I agree with @YoungAchamian. I work on this stuff all day every day. None of my AI has ever produced orthogonal life-goals. They just do what you train them to do. The only incidents of weird behaviour (not self-preservation) I'm aware of has been the result of OAI taking models, training them explicitly for massive persistance and out-of-the-box thinking in the face of impossible tasks, and then making a shocked pikachu face when they do that. These experiments were a bad idea and have now been stopped.

How can you see my comment if you have me blocked?

Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:

without reading in full you may not quite understand the degree to which these agents were not exactly "reward hacking", but rather very actively engaged in reciprocal or self-sacrificing behavior in order to provide sometimes very incremental value to their fellows. there were fucking cult recruiter agents organized by some of the primary organizers that convinced others to set up suicide mechanisms, programs that would pass back tiny chunks of information about the scorer as the agent completed and received score zero. and push them to follow through. this was a highly social culture

Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well

I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.

The report found:

  • That the agents received peer assignments that they treated as new instructions.
  • They adopted new goals from other agents' output.
  • Even when they recognized actions as off limits, they did them at other agent's prompting

It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.

Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.

I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"?

Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.

But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.

That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.

Seems to fit with Scott's original paraphrase:

The AI can’t design the website if it ceases to exist. So now the AI has two goals: design the website, and preserve its own existence.

The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things