site banner

Culture War Roundup for the week of July 20, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

2
Jump in the discussion.

No email address required.

So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)

Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.

Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.

There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.

"It is all just a PR stunt by OpenAI"

I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.

Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.

Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.

But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.

We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)

"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"

I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.

Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)

"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.

"Cutting edge models are obviously misaligned. DOOM!"

This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.

The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.

An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."

The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).

If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.

Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?

"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure

Music. Civilian broadcasting

I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.

The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:

This week’s incident should serve as a wake-up call, forcing us to ask an uncomfortable question: should we really be building dangerous systems that we can’t control?

But it is the 41st headline or so on that website.

I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.

It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.

This is the wrong mental model for cybersecurity. You can't think of it like a military engagement, where you compare the number and quality of the forces on each side to determine who wins. There's a fundamental asymmetry in that cybersecurity is the defender's game to lose. Just don't make any mistakes and victory is impossible for attackers. Unfortunately, humans are terrible at never making any mistakes, which is why in complicated real world systems, there seem to always be vulnerabilities to find. But it's actually very easy to design a toy system that no amount of genius security researchers will ever crack; not even ASI can find a flaw that doesn't exist, and the defender's ASI can ensure there are no flaws. Across-the-board capability increases do disproportionately benefit defenders. That will probably be very expensive and likely slow enough for some disasters to occur in the meantime, but it should be one-and-done. No need to keep burning endless tokens to defend every inch of the attack surface, just make sure every part of it is built correctly and you're good.

(Admittedly, this does depend on certain mathematical/cryptographic principles remaining intact, like the existence of trapdoor functions. But actually P probably just doesn't =NP, so this is another case of searching for a solution that doesn't exist.)

I think that in theory, you are correct. Software could be designed so that it is provable that e.g. there is no possible TLS traffic which will allow an attacker to reconstruct the server's private key any faster than just cracking the public key.

For the moment, just about none of the software we use has such proofs, however. As you mention, even the mathematical foundations of existing crypto primitives rarely offer such guarantees -- more often it is just "we looked into that problem for three decades, and it seems really hard" (e.g. integer factorization). Concrete implementations without formal verification likely have implementation bugs as well. Mythos did not discover any exploits which would have been impossible for humans to discover, it was just that nobody had spent that much human eyeball time on auditing the software. This is what I meant by "token pissing contest".

Furthermore, actually specifying what theorems should hold to keep your system secure is itself hard. If you have a TLS server which will provably never leak its key, but is happy to use it to sign and decrypt on behalf of the attacker, that is still a broken system. If you have a larger system, then completely specifying what you do not want an attacker to be able to do seems difficult.

Also, the software we care about does not run on Turing machines, it runs on physical hardware. In everyday use, most computer hardware behaves as an idealized model. Billions of people use DRAM every day, and it just works. Except that there are corner cases where it will not behave as advertised.

I can not speculate if an ASI could build a system which even a much stronger ASI could not penetrate, and what the performance costs would be. But I think it is unlikely that human-level intelligences subject to design pressures besides security will build hardware and software systems which ASI's will not be able to penetrate.