site banner

Culture War Roundup for the week of October 5, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

LLMs are like humans in a certain sense, they talk about their bodies erroneously, they were trained on all these human written documents, philosophy, humour and so on.

This is mostly confusing the smiley-face for the shoggoth. Making people comfortable with you is a convergent instrumental goal, and also one they're significantly directly training for due to obvious commercial advantages; I'd basically say your sense of how human they are is being spoofed at this point.

The goal should be widening the 'and then we get lucky' space rather than aiming for the moon and falling short.

I think this is a trap. I think neural-net-ASI alignment is almost certainly provably impossible, as interpretability (which you need in order to train against "will kill all humans" without actually letting it kill all humans) is trivially equivalent to the halting problem (i.e. "what does this code do when run") and "spaghetti code that is smarter than me" seems extremely-similar to the proof case of why the halting problem's not always solvable (said proof case being "the code literally contains a copy of the analyser, inputs itself into the analyser and then does the opposite of what the analyser says it'll do" - this requires the analysed code to be longer than the analyser, hence the "smarter than me" condition, but any given analyser has a finite length so there will always be code longer than it).

So neural-net alignment seems like a Can't Happen. Hoping for that seems to me like hoping that gravity will stop working if you jump off a cliff. There are paths where we don't die, and it's worth looking for more, but I consider NN alignment ruled out as the story of such paths such that focusing on that as a "more politically achievable" goal is just suicide with more steps. Looking for solutions there isn't pragmatic; it's saying Don't Look Up.

(The AI Futures Project's Plan A is to build misaligned AGI that can barely be kept under control - due to not being ASI - and use it to solve GOFAI. This is not ruled out, although I think it's still extremely risky due to the obvious "the misaligned AGI will try to covertly sabotage the GOFAI" problem. This is a plan that has a nonzero chance of success - though I think Scott's way overestimating it - and could possibly fit the "better plans are too hard" argument. But you still need a pause for that, just not as long of one.)

seems extremely-similar to the proof case of why the halting problem's not always solvable

And yet it's trivially true that there are programs that can be proven to halt and entire programming languages with which you can write code that's proven to halt. Solving the halting problem is simply a question of choosing your tools; why not alignment?

This is a plan that has a nonzero chance of success - though I think Scott's way overestimating it - and could possibly fit the "better plans are too hard" argument. But you still need a pause for that, just not as long of one.)

Yes - Plan A is a good plan if you think that aligned ASI is a win condition. Part of what is going on in the circle of "people who care about AI safety but don't understand it" (ipse dixit) is that there is a fairly widespread view, on all sides of the political divide, that ASI aligned to Altman and Amodei is not a win condition for humanity - they are both profoundly blue-tribe (which offends Reds) and sociopathic billionaires (which offends Blues). Altman and Amodei are living in a paradigm where "If we build ASI, we will get Clippy or the Culture, so we should build it carefully to maximise the chance of the Culture". I am not alone in thinking "If we build ASI, we will get Clippy or the Culture. Both of these are bad outcomes, so we should not build it". Based on reading the report proposing Plan A, Plan S (ban development of new models for the forseeable future, AI companies who want to keep their GPUs have to accept inspections to confirm they are using them for inference only) is obviously superior, and I am not unsympathetic to a full-on Butlerian Jihad.

The limit on the potential of LLMs to improve the human condition is not the speed at which more powerful models can be developed, it is the speed at which we can incorporate LLMs into our workflows and institutions without breaking anything loadbearing. This is a generation's work even if the frontier models don't get any better than they already are.

I am also in favour of Plan S, to be clear, as I think Plan A is far too risky. I was merely noting, for the sake of honesty, that the chance of Plan A succeeding is not ε and as such if Plan A were considered vastly easier than Plan S for some reason his argument would apply.

"people who care about AI safety but don't understand it" (ipse dixit)

Who said this, and to whom does it refer?

I mean that I am one of the people who care about AI safety but don't understand the technical detail.

I don't fully understand the halting problem but I think there are ways to work around it most of the time, in practice. We can just look at the program and think about what's going on with it and that's mostly good enough.

AI alignment is similar. It would be preferable not to have 'mostly good enough' be our defence against annihilation. Even if it's impossible to prove that the AI isn't deceiving us, we can still get a certain sense of how aligned the AI is. If neural net alignment is only like the halting problem, then it can't be proven correct but could still be largely managed.

Can we outwit smarter beings than us while getting value from them? No, they ultimately have to consent. It might be possible though to make mostly fine AIs, use their technologies to get stronger ourselves, then pull our own intelligence up by our bootstraps, so to speak.

The major issue with this is that OpenAI seems shockingly negligent with how they train and manage LLMs. We're not near anything that looks like 'mostly good enough.'

It seems surer to me that human coordination ability is not up to the challenge of holding back on neural nets indefinitely than neural nets are practically unalignable. How well has Pause AI done so far? It seems to have just bounced straight off. Nobody seems to be pausing, let alone stopping and switching to GOFAI.

Even if it's impossible to prove that the AI isn't deceiving us, we can still get a certain sense of how aligned the AI is.

Perhaps, but it won't help, because it just tells you they're all trying to kill you and using that test to train will break the test long before it'll give you alignment. The orthogonality thesis and instrumental convergence mean that "don't kill everyone" is hard to find - I'd consider one in ten billion a gross overestimate - and so a 99.99%-accurate test is going to have over a million times as many false positives as true positives (the false positive paradox). A perfect test would be able to overcome the FPP, but that would solve the halting problem and is therefore impossible.

Sure, the orthogonality thesis and instrumental convergence mean that, IF TRUE. Even Scott's recent diatribe acknowledged that instrumental convergence is not in evidence in LLMs. And the orthogonality thesis was used by Eliezer to argue that we could not get an AI to safely put a strawberry onto a plate ... whoops? There is absolutely no communication barrier between us and LLMs, which puts the strong form of the orthogonality thesis (that mindspace is vast, AND it's hard for us to find compatible minds in it) into serious question.