site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.

There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.

"It is all just a PR stunt by OpenAI"

I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.

Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.

We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)

"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"

I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.

Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)

"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.

"Cutting edge models are obviously misaligned. DOOM!"

This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.

Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?

"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure

Music. Civilian broadcasting

I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.

The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:

This week’s incident should serve as a wake-up call, forcing us to ask an uncomfortable question: should we really be building dangerous systems that we can’t control?

But it is the 41st headline or so on that website.

I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

Isn't this a bit premature? If the prompter said something like "use all avenues" or "this is critical to save the life of my mother" and the model's safeguards - which presumably instruct against such behavior - were turned off (as OpenAI says they were) then it seems quite possible that the AI was performing as instructed.

Air-gapped systems are a PITA to run

Just remove the Wi-Fi antennae.

Air-gapped systems are a PITA to run

Just remove the Wi-Fi antennae.

The computer running the evaluation didn't have internet access.

The first step in its hack was breaking out of its sandbox to take control of its computer. The second was hacking the OpenAI internal network until it found the internet. The third was hacking HuggingFace.

A proper airgapped computer couldn't access anything off of its own hardware. As a random example, it couldn't receive data from an LLM running in an off-site data center, which would make evaluations difficult.

A proper airgapped computer couldn't access anything off of its own hardware.

Sure, but I think for these purposes you need the network to be air-gapped while you are actively running tests. (Unless you expect your model to be writing malware that you can't detect that activates when it is not running, in which case you should be air-gapped anyway.) So you can plug the Wi-fi in for things like updating software between test sessions. Which makes air-gapping a lot less painful.

You're right that you'd need to run inference locally, though. But I don't actually think it's remotely beyond OpenAI's capabilities to stand up enough local compute to run a few test instances of a frontier model.

Sure, but I think for these purposes you need the network to be air-gapped while you are actively running tests. (Unless you expect your model to be writing malware that you can't detect that activates when it is not running, in which case you should be air-gapped anyway.)

Would you have recommended (your easy, updateable) air-gapping before you saw this failure?

Do you predict you would recommend (hard, strict, one-way) air-gapping before models write non-detectable autonomous malware that could escape your soft airgap?

OpenAI didn't take the unknown risks seriously enough to pay the high cost of airgapping the test computer, and therefore they didn't take security seriously enough to prevent the attack. I suspect that this general attitude will carry forward, and corrections will only happen after failures. Next time might be more serious than a benign attack on a friendly company.

I don't actually think it's remotely beyond OpenAI's capabilities to stand up enough local compute to run a few test instances of a frontier model.

Sounds like a PITA, and much worse than "just remove the wifi antannae".

Would you have recommended (your easy, updateable) air-gapping before you saw this failure?

No, my past recommendations have been "plant a nuke under the datacenter." More seriously, it probably would have depended on the characteristics of the models behavior in the past, and the characteristics of their setup, which I am not privy to. It sounds like OpenAI had good reason to believe that their setup was not vulnerable.

Do you predict you would recommend (hard, strict, one-way) air-gapping before models write non-detectable autonomous malware that could escape your soft airgap?

Well, I could recommend it now and then trivially answer "yes." And the actual answer to that question is more "well can the model write malware without you people who wrote the model and have access to the software and hardware on the machine noticing?" But apparently they don't monitor their models during testing well enough to notice an involved hack-a-thon (understandably, watching a model grind inference is boring) so the answer to that might be "no, even if it is physically possible we're not necessarily going to take those steps."

Let's say that based on what I do know, I think it would probably be a good idea, and if they don't take this step they should take others. Even if you aren't worried about existential risk, good old reputational risk and legal exposure I think is a good enough reason to do this.

Sounds like a PITA, and much worse than "just remove the wifi antannae".

You're right that I probably should have taken the storage-for-weights-and-cooling problem a bit more seriously (although I guess if this model is actually very optimized then it might be able to run on a local machine, which would make it pretty easy). So I concede that setting it up properly would be a PITA. But once you set it up, I don't think actually running it would be hard ("remove the Wi-fi antenna"). They definitely have enough money to build that, or buy a welding workshop somewhere with the electrical infrastructure, server space, and possibly cooling requirements already in place and air-gap it. So it's hardly a PITA for a company that is already working to build dedicated datacenters, in the big scheme of things.