site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.

There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.

"It is all just a PR stunt by OpenAI"

I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.

Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.

We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)

"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"

I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.

Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)

"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.

"Cutting edge models are obviously misaligned. DOOM!"

This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.

Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?

"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure

Music. Civilian broadcasting

I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.

The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:

This week’s incident should serve as a wake-up call, forcing us to ask an uncomfortable question: should we really be building dangerous systems that we can’t control?

But it is the 41st headline or so on that website.

I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

Isn't this a bit premature? If the prompter said something like "use all avenues" or "this is critical to save the life of my mother" and the model's safeguards - which presumably instruct against such behavior - were turned off (as OpenAI says they were) then it seems quite possible that the AI was performing as instructed.

Air-gapped systems are a PITA to run

Just remove the Wi-Fi antennae.

Air-gapped systems are a PITA to run

Just remove the Wi-Fi antennae.

The computer running the evaluation didn't have internet access.

The first step in its hack was breaking out of its sandbox to take control of its computer. The second was hacking the OpenAI internal network until it found the internet. The third was hacking HuggingFace.

A proper airgapped computer couldn't access anything off of its own hardware. As a random example, it couldn't receive data from an LLM running in an off-site data center, which would make evaluations difficult.

didn't have internet access

still has access to the internal network

Bruh.

There was a twitter post a while back that basically went; 'Oh, you think you're so smart and know more than the experts?!' 'No, I think I'm a fucking idiot and know more than the experts. That's the problem.'

People should just hire me as an AI safety expert, because even I know that 'access to an internal network' is basically just another way to spell 'access to a bug-ridden security nightmare'.

Grant_us_eyes: "Hey, this is clearly insecure and dangerous. If we want to give an untrusted agent full range to do whatever it wants to observe its capabilities, we need to build an airgapped system with no network access. All new evals, code, and weights will need to be transferred to it by USB stick. And only USB sticks we've fully audited for exploitable firmware."

OpenAI: "Uh, won't this slow us down? kthxbye"