site banner

Culture War Roundup for the week of September 7, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

5
Jump in the discussion.

No email address required.

AI X-Risk Goes Mainstream?

Jacob Coxon:

"I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below."

116,400,000 views, 634,000 likes.

Evan Hubinger (head of alignment at Anthropic):

Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.

34,000,000 views, 51,000 likes.

None of this is substantially different from what certain AI industry insiders have been saying for 5-ish years now, but it's clear from the replies, quote-tweets, and engagement statistics that most people are hearing about this for the first time. Yes, there was clearly some coordination with friendly media prior to the announcement given the almost simultaneous WSJ article (whereas news outlets were initially reluctant to publish anything on the unexpected news of OpenAI solving a Millennium Prize Problem). One can take the maximally conspiratorial position that this is a 4D bankshot by Democrats to end tech innovation forever, but it was going to break through eventually.

In other words, Coffeezilla is now on the case.

I agree with @Shakes that there's a good chance that this is all just for show to build up hype around Anthropic, just like... basically everything they do. But if it's not, then we should throw every single person involved in prison for the rest of their lives. If they actually believe that AI has a 10% chance of killing us all, and research it anyways, then their willingness to roll those dice is a threat to the rest of us which we shouldn't ignore. There basically are no "oh shit we shouldn't do that" limits that such a person will ever respect.

The "it's all just marketing" people need to reckon with the fact that they were rushing to declare the same thing over the Hugging Face incidents, only for the German wiki revelations to completely destroy that argument.

That was a different company obviously, and it doesn't preclude future incidents from being leaked for marketing purposes, but it should still lead to a serious adjustment in beliefs on what is and isn't marketing

Well, the thing with the marketing arguments is that there's many ways to make those arguments, some more plausible than others.

Strong form: OpenAI intentionally induced their agents to compromise the internet for marketing purposes.

I agree that trying this would be very stupid and not particularly likely.

--

Semi-strong: OpenAI intentionally sandboxed their agents poorly to stochastically induce agent misbehaviour for marketing purposes.

This one I could see going either way. It does seem like there were some egregious oversights in the sandboxing setup, but whether these were intentional or real oversights is impossible to say.

--

Weak: OpenAI did not intend for misbehaviour to happen, but now that it has happened they're spinning it as hard as possible for marketing.

Personally, I think this one is pretty likely. I was very unimpressed by the METR report for instance, they basically just slopped together some agent review to hype up capabilities without actually auditing the root causes of why it all happened in the first place.

The problem with the weak argument is that it is, well, weak. It demonstrates nothing, particularly given we can already assume that every corporate communication ever made is being spun to portray the corp in a more positive light.

Indeed, we actually know one of the methods that OpenAI used to try and spin the METR evaluation more positively: they heavily restricted the scope to prevent any of the details of the German wiki collusion coming to light. Which also raises the other critical weakness of the marketing argument, in that many of the recent safety scandals have come to light from third party sources. METR, the UK's AISI, and a couple of independent researchers for the collusion stuff. Now, given METR's links with the wider AI ecosystem, it is not too much of a stretch to suggest that they could have been influenced to present things in a certain way and thus aren't fully independent, but it would be a much larger leap to suggest this of a UK gov organisation, and an impossible leap to suggest that of a couple of randos.