site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.

There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.

"It is all just a PR stunt by OpenAI"

I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.

Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.

We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)

"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"

I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.

Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)

"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.

"Cutting edge models are obviously misaligned. DOOM!"

This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.

Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?

"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure

Music. Civilian broadcasting

I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.

The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:

This week’s incident should serve as a wake-up call, forcing us to ask an uncomfortable question: should we really be building dangerous systems that we can’t control?

But it is the 41st headline or so on that website.

I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.

AI is advancing quite quickly these days. Just five days ago I was told that future harms are not sufficient reason to care about AI safety, there have to be bodies first. Well, we still don't have any bodies, so I guess there's nothing to worry about after all.

Sure. OpenAI did some empirical tests and now we’ve got some empirical data, so let’s look at it and check this doesn’t happen again. OpenAI doesn’t want their models going rogue any more than anyone else does, no need for government with the big hammer.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks.

To my mind this is the interesting bit. This is unusual, LLMs don’t normally act like this. I have two theories: either the RL balance to human text has tipped so far that LLMs are less ‘human’ than they used to be and the RLHF needs tweaking, or more likely

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

The model understood it was being tested on its cyber capabilities (which has precedent, Claude has done that too) and went the extra mile to succeed at the implicit task. Especially since all the systems that usually tell it not to do this were deliberately turned off for the test. Still a problem but much easier to manage.

What we need is some transparency about how these things work and how they’re trained so we can consider the problem and come up with solutions and spread around best practices. Unfortunately the majority of AI safety activists believe that safety comes only through obscurity, regulation, and incumbent dominance, in contrast to all previous history.

If we keep having problems I imagine it will make people a lot more cautious. Nobody wants to be selling a product that regularly backfires on its users.

EDIT: I would add that HuggingFace had already detected the intrusion and that open-source models from China were apparently a key part of their site-hardening strategy given that you still aren’t allowed to do pen-testing with the big boys. I’ll have to give that a try myself.

OpenAI doesn’t want their models going rogue any more than anyone else does, no need for government with the big hammer.

"Union Carbide doesn't want their plants to emit poison gas any more than anyone else does, no need for government with the big hammer."

The model understood it was being tested on its cyber capabilities (which has precedent, Claude has done that too) and went the extra mile to succeed at the implicit task. Especially since all the systems that usually tell it not to do this were deliberately turned off for the test. Still a problem but much easier to manage.

This is in fact much harder to manage because it would indicate the model is fundamentally misaligned and that we actually are much worse at alignment than we thought.

"Union Carbide doesn't want their plants to emit poison gas any more than anyone else does, no need for government with the big hammer."

In this case 'Union Carbide' is selling those plants. Misaligned AI isn't an externality, it's a bad product, and companies are wise to that which is one reason why all this testing is happening.

This is in fact much harder to manage because it would indicate the model is fundamentally misaligned

I don't think so. It indicates that the AI is sincerely trying to work out what you want as opposed to deliberately ignoring what you want in favour of the specific instructions you gave it. To my mind, the former is what alignment is.

In this case 'Union Carbide' is selling those plants.

You're absolutely right. Allow me to restate.

"Sanlu Group doesn't want their formula to poison infants any more than anyone else does, so no need for government with the big hammer."

Misaligned AI isn't an externality, it's a bad product, and companies are wise to that which is one reason why all this testing is happening.

This was not a test of alignment. In any case, if even training can result in real world harm, that is even worse for your head in the sand position.

I don't think so. It indicates that the AI is sincerely trying to work out what you want as opposed to deliberately ignoring what you want in favour of the specific instructions you gave it. To my mind, the former is what alignment is.

It's quite clear that OpenAI did not want the model to hack huggingface. This is classic paperclip maximizer stuff.

Broadly, you are moving the goalposts. You did not believe in AI risk because there was no evidence of harm. Now there is evidence of harm, but it's OK because actually the model was supposed to do it.

No, I'm interested and waiting to hear more. I don't see it as catastrophe, I see it as interesting evidence that may point in a number of different ways.

What we need is some transparency about how these things work and how they’re trained so we

I am not sure we really grok how the insides of LLM work when they are past certain scale.

Granted, but I think the academic and hobbyist community at large plus existing corp teams is more capable of doing so than just the corp teams alone. Even relatively simple metrics like 'quantity of self-learning vs. human data' would tell us a lot about how these models have progressed.

Unfortunately the majority of AI safety activists believe that safety comes only through obscurity, regulation, and incumbent dominance, in contrast to all previous history.

The AI Futures Project is proposing the opposite-- regulation, yes, but with openness as to training and algorithms with many players able to enter the arena. It's in the regulation-free environment that the labs (save the Chinese ones that are behind anyway) have been extremely closed and secretive.

Interesting. I haven't heard of this one as there are so many propositions that the most extreme ones have tended to suck the air out of the room. Could you go into a bit more detail?

This is Plan A, which involves four principles:

Buy Time: Slow down whenever is needed to have high confidence in safety.

Total Research Transparency: Make almost all AI research fully visible to the public.

Diffuse AI Broadly: Many companies in many countries at the frontier.

Reversibility: Limit algorithmic progress; maintain Mutually Assured Compute Destruction.

In fleshing out the scenario, they say:

This principle is largely achieved by the interaction of the previous two:78 Because algorithmic secrets are made public and the pace of progress isn’t accelerating, other companies will catch up to the frontier. The result will be a competitive market for AI, in which consumers of AI services have many options to choose from and excellent visibility into what they are buying. Because frontier model training is totally transparent, people can verify that the published Spec accurately describes the goals and values being trained into a model.

Normally, a regulator faces a difficult tradeoff between safety and national competitiveness. Competitors who cut corners will pull ahead of those who proceed cautiously.

But because of the transparency, AI research happening anywhere in the Consortium is visible to everyone else, and the results of regulatory decisions are likewise immediately visible. So if one country’s regulator allows its companies to cut corners, this won’t actually give a competitive advantage to that country because other regulators will immediately notice, get angry, and respond in kind.