site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

You invoked sci-fi

I'm sorry but this seems like an uncharitable assertion in the extreme. I'm not the one that defined the problems with AI in relationship to Malevolent AI Gods, Singularities, Outcome Pumps, Roko's Basilisk, Grey Goo/Paperclip maximizers, Magical Genies, Shoggoths, Goodhart demon, etc. These are all inherently Sci-fi ideas. I have no idea how you seem to think any of these relate to how Neural Networks, or Bayesian Networks, or any other of the myriad ideas ML/AI engineers have designed.

I did and it remains linked.

I misspoke, I don't tend to read links other than to fact check them as referencing the quote. If you want to quote a bunch of sections to make an argument, please do. But yeah I'm not going to an external blog/substack, if it's important enough to make your argument, its important enough to copy and paste on the motte.

The argument is made at length in the piece. And any many other places, it straight credulity that you have not seen it.

There is a lot of things to need to be read out on the internet, many of far more actual importance than some wannabee philosophers who really like typing out long screeds, why should I dig through thousands of blog posts? If the argument is important enough, I'm sure the adherents can post the good arguments for me, summarizing what the actual position is.

The argument is simple:

#1: Sure, AIs get more capable all the time, how do you know there isn't an asymptote?

#2: How? By what mechanism? There is a lot of assumptions about capabilities baked into this deceptively simple axiom

#3: Nice back to Sci-Fi Goobly-gook. Plainly stated: We cannot guarantee that an AI will correctly infer what a user really wants, avoid collateral harm, and act in everyone’s interests. Congratulations you've converged on a fundamental question for any human system, replace "AI" with a "human" and the sentence is trivially true of anyone in any system.

#4: conceivability is not evidence of inevitability, a familiar narrative is not a causal series of events.

#5: More Sci-Fi, I thought I was tilting at Sci-Fi windmills? Literally: "Once the AI becomes capable enough, it becomes the machine god and humans are included in its omnipotent calculations in a way humanity might not survive"

Then we should not build it.

Again this assumes that this eventually AI will be sentient and that by building a sentient AI we will create a "alignment" problem. This is an extraordinary claim requiring extraordinary amounts of evidence. So Prove it.

I thought the sneer you had was that alignment people were ridiculously thinking they were sentient? It seems like something you believe more than them.

I can decouple my disagreement with the word "misalignment" with the actual argument being thrust forward by the term. I'm stepping into the frame of Yuddites, to point out how even in their ontology it is a non-solvable problem. Not only is a non-solvable problem, its not a new problem, meaning it does not need to appropriate some new word to describe it.

My argument for what "misalignment" actually is. Take the LLM-Agent we have today, it's a complex system, it makes errors and apparently there are no controls in the system to account for those errors. How does this system work on an engineering level?

  1. Large Transformer Model (LLM) is trained on predicting the next token in a series on basically the whole of human written data, both analog and digitally.
  2. LLM is then trained via RL variations to produce responses that human evaluators prefer, there is some level of directionally correctness in this. Sounding non-human is often because it sounds wrong, but the r-value is not 1.0. Some safeguards are trained into the model here.
  3. These days we supplement #1 with training on code/math data and additional task decomposition training. The goal is to get the model to be able to logically decompose a bigger problem into smaller problems and then solve those smaller problems
  4. An API harness is applied to the model this filters out some prompts, adds system prompts to the model and a whole host of other things. This is just a software layer. This is what 99% of people interact with.
  5. Agentic AI Harness is applied. This is a software program that recursively prompts the LLM and feeds its output back into it, in order to accomplish decisions, it also takes LLM outputs and executes the code it is given.

So when a "misalignment" occurs what happens? Well hallucinations are really an error with #2, the model is not designed to be "correct" it's the r-value problem. It looks right but isn't. It produces tokens that are correct in the next sequence, it reads like what a human wants to see, but it's not true. There is not intent. Given its learned distribution and current context, the model just generated a highly plausible continuation that did not correspond to reality.

The model does something it shouldn't? #3 is the problem, it output the incorrect task decomposition, and since the system is automated with no guard rails it just executes that task decomp. And sometimes its even #4 and #5 having errors thrown in. Turns out complicated systems of algorithms behavior in logically consistent and coherent ways that someone forgot to error check. We don't scream that "software algorithms are misaligned" because some coder forget an unit-test on an edge case. These aren't "misalignments" they are are classic errors on a non-perfect model in a system lacking in classic control theory.

Believers in some version of their cause are currently littered throughout the frontier labs. You have an impossible standard here where you blame them if they're on the frontier, for clearly not believing in what they preach, or blame them for not being on the frontier as then they must be uninformed cranks. No way to win, you never have to actually think about it. But yes, of course they don't have an alignment mechanism, their whole point is that the problem is incredibly hard and that we have no solved it, that we might need to spend decades solving it, but that the alternative is everyone dying.

Sure and there are christians, mormons and muslims in the physical sciences. I don't begrudge people their religious beliefs. If a frontier AI researcher wishes to belief in alignment-problems that his/her/their belief. The Bay is quite literally for Rationalists like Utah is for Mormons. I can still think it's a silly belief and I still think the word misalignment has been invented to describe a problem with system error by a bunch of non-engineers. And the lack of actually solving the problem, or even making progress on it, is indicative of a general grift specifically for something like MIRI.

You can use the word "sneer" to describe my behavior, but what word would you use if you wanted to point out holier than though attitudes among some christian suicide cult? Acting ridiculously gets you ridicule, it's not "sneering".

I'm not the one that defined the problems with AI in relationship to Malevolent AI Gods, Singularities, Outcome Pumps, Roko's Basilisk, Grey Goo/Paperclip maximizers, Magical Genies, Shoggoths, Goodhart demon, etc. These are all inherently Sci-fi ideas. I have no idea how you seem to think any of these relate to how Neural Networks, or Bayesian Networks, or any other of the myriad ideas ML/AI engineers have designed.

These are all intuition pumps and metaphors for explaining the actual intuition. Goodhart isn't a science fiction concept, he explained that under optimization pressures a system that satisfies your stated objective can fail your intended objective and in fact the better the system gets at optimization the better it gets at trading off your intended objective in pursuit of better achieving your stated objective. This thread started out because of an example of this very principle playing out in reality.

Again this assumes that this eventually AI will be sentient

No. And this is the third time it's been explained to you so let me be extra blunt. No part of alignment requires or cares about sentience, wants or experiences for the system. The lights need not be on. You didn't even have to click into the short essay, the paragraph I posted covers this.

These aren't "misalignments" they are are classic errors on a non-perfect model in a system lacking in classic control theory.

Error implies the quantity shrinks as you train it more, this is true of hallucinations, they've been reduced. A hallucination doesn't help you achieve a target goal. This is not true of misalignment problems because misalignment is reaching the goal through unintended and harmful pathways.

The coastrunner boat spinning infinitely scoring forever isn't failing to minimize loss. It's minimizing loss perfectly. But the loss target was not aligned with what we actually want.

Your #2/#3 ignores a failure mode monotonically worsened by the thing that fixes every other failure mode, which is why you keep having to reach for "someone forgot a unit test." All the unit tests pass, it's performing exactly as expected. You can't expect to build smarter and smarter scaffolding to swatt it away from optimizing around your guard rails. Eventually it gets strong enough and you lose, game over.

We don't scream that "software algorithms are misaligned" because some coder forget an unit-test on an edge case.

Grep doesn't go looking for novel ways to satisfy your query, it's simple executing a list of instructions. Deterministic software fails when it runs into states you didn't specify at compile time. Optimizers fail when where you specified correctly but it turns out the specification wasn't what you actually wanted. They're a different class of problem.

#1: Sure, AIs get more capable all the time, how do you know there isn't an asymptote?

Obviously we don't and I'm pretty careful to specify, as I did in the previous comment, that this is one of two valid outs. If you think these systems won't get much more capable then sure, although I'd like you to demonstrate that because people predicting it have been very wrong for a while now. But if that's your contention then I think your whole argument would be much different. The whole ask of the AI safety crowd is essentially to impose an artificial asymptote.

#3: Congratulations you've converged on a fundamental question for any human system

Yes. And we have not solved this problem with humans either although we've spent a tremendous amount of effort attempting to better mitigate the problem. contract law, courts, fiduciary duty, audit, separation of powers, criminal enforcement, a state monopoly on violence. Half the discussions on this forum boil down to disagreements on how exactly we're failing to solve this problem. And that's between humans who are individually squishy, mortal difficult to scale(see corporations and their shenanigans in the 17th century for what happened immediately after we found novel ways to scale human effort), must sleep and have evolutionarily built in game theory instincts.

Sure and there are christians, mormons and muslims in the physical sciences. I don't begrudge people their religious beliefs. If a frontier AI researcher wishes to belief in alignment-problems that his/her/their belief. The Bay is quite literally for Rationalists like Utah is for Mormons. I can still think it's a silly belief and I still think the word misalignment has been invented to describe a problem with system error by a bunch of non-engineers. And the lack of actually solving the problem, or even making progress on it, is indicative of a general grift specifically for something like MIRI.

This proves far too much. It labels all speculation of dangerous outcomes as fundamentally religious. Is climate change fundamentally religious? Fear of demographic replacement? Pick your conservative or liberal fear. This position dooms to you dying out to the first large scale novel problem you encounter.

And I dispute heavily that it's a bunch of non-engineers that are concerned. Some of the biggest voices in the movement are the frontier engineers including the CEO of the lead frontier lab. If you're going to dismiss anything the non-engineers say as not credible because it comes from non-engineers and what the frontier engineers say as non-credible because it's "hype" then you've made your position unfalsifiable.