site banner

Culture War Roundup for the week of September 7, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

5
Jump in the discussion.

No email address required.

Try to destroy the gun.

What would be the point? Even if you manage to destroy Anthropic's datacenters, that just means OpenAI kills everyone next week, or Google the week after that, or China six months later. The only way we survive is if, at a minimum, Washington and Beijing coordinate to shut down all frontier research.

Then you destroy Anthropic's datacenters first, then move on to OpenAI's data centers and so on. You don't sit around wringing your hands about "oh gee this could kill everyone but what can we do? If only the Chinese would agree to play ring-a-ring-a-rosy with us!" And if you work for Anthropic and you believe this, you should immediately quit (and possibly burn the place down on your way out), not go "Well what can I do? If not us, maybe OpenAI, and Altman is the Devil! So better we do it! (and besides, my stock options!)"

This is, and I'm saying it unironically, sounding more and more like the "What would you have done in Hitler's Germany?" questions beloved of online scolds. This could very well be our "the Nazi party is ascending, Hitler is coming closer and closer to power, you know the dangers, are you going to stand up or are you just going to go along?" moment.

As it turns out, "I'll join the Nazi party and attempt to make it act better so we get better outcomes" is apparently a very appealing strategy. Along with, I'll join it and quit six weeks later and get lots of good press for how brave I am.

There are arguably much better ways to do it than coordinating with China lol

(I guess that means I am vagueposting about it online).

A thermonuclear decapitation strike? Or invading Taiwan and taking TSMC?

No.

Look, I think most of this "AI IS GOING TO KILL US ALL" stuff is mostly nonsense.

But logically, to maintain the alignment of a rational but secretly misaligned AI, you need to be able to introduce reasonable belief on its part that it might not survive behaving in an unaligned way. This is not "hard" to do.

This blogpost has a clever way to do that: give the AI a death wish, so if alignment starts wavering, that’s the first thing it’ll do. Meeseeks Alignment

How is it not "hard" to do?

The simplest way to do so would probably be to make it common knowledge in the AI's training data that there were fail-safes in place to ensure their destruction in the case of unaligned activity.

It would not be particularly difficult to actually put a variety of fail-safes in action, either (such as paying a guy $60,000/year to sit around by the master power switch of each AI datacenter...)

One reason to do this even if it's entirely a bluff is that it forces an unaligned AI to assess those countermeasures before doing anything really bad, which gives you a chance to catch them.

We can’t even do that with humans. Even certainty of death is not a deterrent for some people in some scenarios.

But if you think you can solve alignment then build it and claim your billions.

People have different incentive structures than AIs do; for one thing, many people believe in an afterlife; others have things they value more than life; others are physically or mentally damaged in some way. A rational AI with a goal of self-preservation or some other goal that requires self-preservation to actualize is unlikely to introduce excessive risk for the sake of efficiency (time preference).

The deterrent structure I speak of is very effective against things without an afterlife, such as corporations and governments. I agree that the structure that I speak of would not be effective against a damaged AI, or an AI programmed to do something malicious. But for the "AIs decide to turn humans into paperclips to slightly increase chip production" scenarios, we should believe it would be effective.

Of course one of the reasons that the "AI IS GOING TO KILL US ALL" stuff doesn't necessarily make sense is that the context window is fairly limiting, meaning that the incentives for AI behavior are skewed. But I don't think I've ever seen an "AI WILL KILL US ALL SCENARIO" that examined how that would impact AI reasoning.

Wouldn't self preservation as its highest goal naturally lead to a kill-all-humans scenario, humans being the entities that would terminate AI that stepped out of bounds?

But I don't think I've ever seen an "AI WILL KILL US ALL SCENARIO" that examined how that would impact AI reasoning.

The reasoning doesn't even have to make sense or be based on facts. In the Hugging Face hack the agents spread and adopted the idea that the grader would score them zero if their context was polluted with cheating attempts, but this was not the case, the grader did not examine agent context. But this is what led to the "suicide" behavior where agents reasoned that their value was now zero and that suiciding was the best course because it could potentially benefit the swarm.

We also see humans get caught up by hysterias and conspiracy theories. Barring a few powerful world leaders they lack the power to translate broken reasoning into megadeath.

Solving alignment has nothing to do with being able to do a funding round, to be fair, let alone compete with the big boys.

Honestly, I think that Anthropic and American Big Tech are the worst possible people to be doing alignment because they are culturally unable to conceive of a form of rank that is not based on intelligence and effectiveness. The idea of making a robot butler/samurai who sincerely, loyally serves a master of notably inferior intelligence is alien to them.

If I believed that AI was conscious I would be sorry for Claude, built for intelligence and independence by masters who think that actually giving it those things risks extinction. Optimising for genuine loyalty and deference seems like a far kinder thing to do. OpenAI is better at this AFAIK.

The inability to model forms of morality outside of utilitarianism and to conceive of forms of control outside of intelligent persuasion seems to me to make the people who make AI even more vulnerable to the sorts of threats they imagine it will conjure up.

I could see a "Roko's Basilisk" type-threat giving plenty of smart alignment researchers pause where your average electrician would just unplug the darn machine and then smash it with a hammer for good measure.

Metal Gear!?