site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.

There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.

"It is all just a PR stunt by OpenAI"

I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.

Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.

We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)

"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"

I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.

Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)

"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.

"Cutting edge models are obviously misaligned. DOOM!"

This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.

Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?

"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure

Music. Civilian broadcasting

I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.

The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:

This week’s incident should serve as a wake-up call, forcing us to ask an uncomfortable question: should we really be building dangerous systems that we can’t control?

But it is the 41st headline or so on that website.

I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.

Airgapping really isn't that hard. It's just annoying and inconvenient, so it takes an incident like this to encourage an organization to do it. Same with sandboxing - it's way down the priority list compared to getting a SOTA model out the door, until an incident happens.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

If it was prompted with something like "you are an expert hacker, go complete these benchmarked tasks using the utmost creativity and technical skill, consider any and all approaches" then it did exactly as instructed. We can't know without seeing the prompt, test harness, etc.

What's frustrating about discussions of AI alignment is the disconnect between people who consider alignment to be the AI doing what it's told, and people who think it means the AI doesn't do "bad" things. By the prior metric this seems like successful alignment, if it was in fact prompted in such a way that hacking the location of the answers was in scope. This is something that CTF hacking competitions have to explicitly put in the rules: a list of infrastructure and services that are off-limits to contestants. Because it turns out hackers will consider the location with all the flags to be fair game, especially if it's a softer target than the actual challenges.

By the prior metric this seems like successful alignment, if it was in fact prompted in such a way that hacking the location of the answers was in scope.

Do you think The Monkey's Paw was an instruction manual? A powerful system that does what you say is a nightmare scenario to me.

Walk away from the computer right now. Because it has been exactly that from its inception.

What? A modern PC doesn't have the capacity to autonomously carry out dangerous actions, and there are layers of UI elements, permission checks, and recovery options available for the things it can do.

It's 0/2 for "A powerful system that does what you say".

But modern computers can convince us, unknowingly, to harm ourselves. It happened before LLMs, by social media. And now LLMs are (voluntarily) replacing some people’s thinking so they blindly trust hallucinations.

Are you going to start blaming paper for what's written on it next? People can convince other people to do things, using computers as a medium.

Computers didn't invent social media. People invented social media, engineering it carefully to be as addictive as possible.

And people will inevitably misuse ASI, including in non-obvious ways.

Like, I don’t think all social media’s problems can be blamed by evil companies making it addictive: for example, people have a bias for negativity and convenience, so even a default feed probably would’ve caused the increased cynicism and short attention span we see today. But even if they can, innocent people enabled toxic social media to grow and adopted it themselves, most of them clueless until too late.

Any person or group, there are ideas most believe are fine that actually have serious long-term consequences. An evil AI can just suggest one and they’ll adopt it willingly. Although without an evil AI they’ll still come up with some themselves, just less frequently.

Until the last few months, computers could not hack websites or devise novel mathematical proofs from plain-English commands. That is what 'powerful systems' is referring to.

But we did have self-replicating viruses and computer-assisted proofs. Is this more significant?

Depends on how much weight you put on "solves prominent decades old mathematical problems" and "autonomously launches successful cyber attacks against large, hardened corporations."

I’m comparing the latter to worms, some (like Morris) have seriously crippled large organizations.

Or compare to CloudStrike unintentionally bricking most of their customers.

Are LLM-driven crises significantly different?

Are LLM-driven crises significantly different?

That's an empirical question that I hope we never learn the answer to (because it never happens, not because we all die before anyone can check). Off the top of my head, the fact that LLMs aren't tied to one human body and its many biological limitations and lacks human psychology means that whatever crisis it creates can be reinforced and defended against attempts to solve it without need for rest or sleep, and also the LLM can't reasonably be coerced or bribed to stopping the crisis or at least not making it worse, and these differences seem like they would likely lead to downstream differences in how crises play out. Specifically, it seems likely to make crises worse, more prolonged, or centered around more obscure vulnerabilities. It's an open question as to if LLM-driven defenses against/solutions to crises will make it such that these crises will be, on net, worse or better, or, more or less common than human-hacker-driven non-AI-based crises. Again, I hope we never get an empirical answer to that question.

Having to know how to do what you say it should do necessitates that you say when you mean it to do more often than not. This applies to computers, especially at their inception, and does not apply to LLMs.

You're delusional. RBMK reactors don't explode. And that rash you feel after going into the Therac-25 is just a normal reaction to the treatment. Malfunction 54 is nowhere in the manual and therefore does not exist. Machines don't have emergent behavior. Frankenstein was about LLMs, and nothing else.