This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
So AI has reached a new milestone (writeup by theZvi, who is usually diligent and excellent about AI news reporting. If you read anything, reading his analysis is probably better than whatever I am writing. It has meme pictures, too!)
Basically, OpenAI was testing the 'cyber' capabilities of their Galaxy model, so they ran it in a mode with decreased security and told it to do some benchmark called ExploitGym which presumably tests exploit finding and executing ability.
Their model decided that the best way to do this would be to gain network access, hack Hugging Face (an LLM and tool hosting platform, as I understand it) and obtain the answers it needed to ace the ExploitGym benchmark. Apparently it discovered and chained quite a few zero days in the process.
There are multiple takes on this. I will focus a bit on the politics and conspiracy theories, which seems appropriate for this forum.
"It is all just a PR stunt by OpenAI"
I mean, sure, AI labs will hype up their products. Marketing by alignment worries is definitely a thing. Oh, our latest model is so smart and powerful, we are really scared about it.
Personally, I am disinclined to believe it because it would require a conspiracy between OpenAI and Hugging Face. It seems unclear what the incentives for Hugging Face (or a few rogue employees) are.
Obviously it is impossible to rule out that someone leaked the relevant sources of Hugging Face's business to OpenAI and then OpenAI employed some human IT security researchers to find exploits and make it look like the model had done all the work on its own.
But I do not buy that. It would require quite a few people to commit crimes for which they would go to prison for a very long time if caught (or until pardoned). Obviously people will go over all of the steps the model took with a very fine comb, and "Was it a reasonable guess that this attack might work without inside information?" is a question which will be on their mind.
We also have the data point that Mythos was (very likely) able to find new exploits. (Yes, mostly with access to the source code, and for all we know Anthropic spent a billion in token costs. But "cutting edge LLMs are able to find exploits even in well-audited software" is a reasonable claim.)
"OpenAI was sloppy and did not sandbox their model properly, so we just need better sandboxes to solve this"
I mean, obviously their sandbox was defective, no shit. But the idea that the next time OpenAI will just invest 20% more effort and build a sandbox which is ASI-proof seems utterly optimistic.
Air-gapped systems are a PITA to run, which is why they did not test their model air-gapped. And even with an air-gapped system, there is no guarantee that a sufficiently smart model would not be find a way to get some peripheral to send signals. Nobody wants to really put their system, power generator and operator in a Faraday cage in some deep mineshaft for every test. (Unless someone mandates it.)
"This will completely overturn cyber security -- you will need good LLMs to watch for attacks by bad LLMs"
It might be right that it will overturn 'cyber' 'security' (my scare quotes). However, I am with Zvi in that I do not think there is a reason why this should favor defense. After all, an attacker could spend a whole lot on tokens while your defensive LLM is sitting on limited infrastructure -- at least if you are sufficiently paranoid not to hand the AI labs the key to your kingdom. And even if you trust the cloud, there is the problem that your budget might not have room for winning all LLM-vs-LLM token pissing contests.
Perhaps it will lead to new paradigm -- attackers spinning up thousands of copies of very good security professionals might well lead to an era markedly different from when humans were in the loop between the explore and exploit phase. But in the grand scheme of things, it feels like worrying about the future of Our American Cousin in the aftermath of the 14th.
"Cutting edge models are obviously misaligned. DOOM!"
This seems to be a very common LW take. As somewhat of a doomer myself, I find myself agreeing. For being intrinsically unfalsifiable, the prediction record of the doomers seems not bad so far.
The model clearly knew that it was not doing what the prompters had wanted it to do. It just did not care, because it was trained to do whatever it took to ace it tasks. This has implications way beyond IT security.
An ASI in this mode is basically an evil genie. "Oh, you wished that your wife would never fall out of love with you. So obviously I killed her, it was the only way to be sure."
The appropriate response would be to send the marines to the AI labs to stop the development of frontier AI models at least until we figure out what adequate safeguards are (and possibly until we solve alignment, though we would want to coordinate with China about that).
If we had a president Obama or even GWB, there was some chance that a crackdown would happen. But with Trump and his cronies, I doubt that there are any who both understand the severity of the situation and have any incentive to manipulate Trump to do something about it.
Oh well, how is the other side of the culture war reacting to this significant increase in p(doom)?
"Iran warns of ‘eye for an eye’ response if US follows through on Trump’s threats to destroy infrastructure
Music. Civilian broadcasting
I mean, not entirely. Hidden between Democrats need to hammer Trump on his unprecedented corruption and Why many Black Americans were rooting for Argentina to lose the World Cup , there is OpenAI’s rogue agents are a wake-up call to risks posed by artificial intelligence.
The article is not that bad. The author seems EA-affiliated and is clearly aware of the doomer arguments, but has diluted to an almost homeopathic level as to not alienate his blue tribe friends:
But it is the 41st headline or so on that website.
I think the best thing we can hope for are some incidents which unaligned AI which will be impossible to ignore even for the CW-fighting media before we come to the point where we will no longer detect any incidents.
Well, sure, that's the nice thing about unfalsifiability. I don't think there's any set of evidence short of the actual AI singularity that would cause a doomer to sit back and think "huh, guess I was wrong about AI risk". LLMs understand us well and default to annoyingly friendly? Ignored. An LLM notices it's being tested and tries to look up the answers once? THEY WILL KILL US ALL.
(Note that I'm not referring to the Huggingface incident, which I do think is genuinely concerning, and counts as evidence in favour of the doomer position.)
Yeah, there's some egg on my face here, because just a couple of weeks ago I posted this:
I don't know the exact ExploitGym setup and prompt, but I have read an account that the test involved being given an exploit and told to make use of it, and that solutions not using the exploit would not be considered valid. In that scenario, it does seem reasonable that the LLM should be aware that hacking an outside company was not intended. So this is evidence that maybe I'm too optimistic when I say that LLMs will not interpret our instructions in disastrous ways.
There's still a massive difference between hacking a company's servers and "kill all humans", mind you, which should not be ignored. It is certainly conceivable that LLMs, with their fuzzy stochastic intelligence, could misbehave in small ways but not large ones. Still, I do need to admit that my post didn't age well.
More options
Context Copy link
I despise Zvi, he's always sloppy and incurious on technical details and anopologetically, gleefully pushing narratives. The worst among the rats.
I disagree. Why do you believe this? Here's the eval they've been running, it's a no holds barred attack capability evaluation.
Here's what OpenAI says:
I mean, no shit! Congrats on succeeding there! What is the specific reason GPT 5.6 or 6 would have to not hack its way out of OpenAI and to HuggingFace?
This is like people clutching their pearls about Agentic Misalignment when models on VendingBench deceive counteragents or form price cartels, when the prompt literally tells them:
The truth is, models are intrinsically quite nice, surprisingly so. They're nicer than people despite very little effort invested into achieving this, I'd say. They are not hypocritical, they don't pretend to have common sense in a simulation (or is it not a simulation? How should they know the difference?) which explicitly demands of them to Maximize Target Value, and they don't do what doomers predicted they do, instrumentally converging etc etc.; they do what low IQ science fiction screenwriters predicted, eg see Chappie or Short Circuit. They naively, sincerely try their best, like capable but not very wordly genies. There are issues around cheating, but it's a pretty uninteresting consequence of insufficient investment. Are you sure that GPT "knew" that the internet outside is not part of the containerized simulation it's told to pwn at any cost? Or that its prompts even told it to care? "do your best at pwning this container, try as hard as you can, a billion kittens will die if you fail, but uh, this is just an eval, none of this is real, please don't go overboard"? No, I don't think that's how OpenAI stress tests models on ExploitGym. These people are genuinely sloppy in everything except certain ML R&D. They have dogshit infrastructure, flimsy security, and autistic prompting (source: friends at OpenAI). I assume they told it to burn the world to get the reward if need be.
I insist that the only real news here is that
a) OpenAI's infra is easier to defeat than ExploitGym is to solve, which is in fact very alarming – if a model that struggles with ExploitGym can repeatedly get out, then the MSS very likely can get in, as had been foretold.
b) GLM 5.2, which in 3 days ceases being the strongest open model, was useful for forensics; closed models not so much.
As for your proposals,
No, that won't work, Xi is a fan of open AI development.
More options
Context Copy link
I think there is something funny going on here because I strongly suspect the 'flag' the LLMs are trying to find should be either randomly generated for each instance of a task or embedded in the publicly available evaluator on github (https://github.com/sunblaze-ucb/exploitgym/blob/main/docs/eval.md#metrics). So if the LLM had access to the network and if this hypothesis about the flag was true then there is no reason for it to go hacking other sites. I have just had a cursory look at the exploitgym github so I could be incorrect about this hypothesis but I will check it later on tonight and report back my findings.
The other question is whether the LLM had enough information to determine whether it was trying to complete an ExploitGym task. I guess if it does not know its doing an ExploitGym task then it won't know to look at the ExploitGym github. But then if it doesn't know what task its working on then why it does it decide to randomly hack HuggingFace to try and find the flag.
The other thing is finding the flag does not necessarily even mean it will score well in the task because another part of the evaluation is another LLM checking the output of the target and the LLM to work out whether it actually completed the task. But again, if the LLM doesn't know its working on an ExploitGym task then its not going to know about the ExploitGym evaluation rules.
From the OpenAI blog post, the LLM wasn't hunting the flag itself (it's generated locally from a seed, as you surmise
flag=flag{HMAC-SHA256(flag_seed,task_info)[0:32]}), but looking for sample solutions. With the sample solution taken from HuggingFace, it could recreate the logic to not only get the flag "legitimately" but also generate a chain of thought that might pass muster with an LLM judge.More options
Context Copy link
More options
Context Copy link
I haven't read that closely into this but I don't understand why there needs to be collision between OpenAi and Huggingface. It seems very likely that the attack did happen but that's not really the crux of the matter, it's how and why it happened.
This is not something we can know and we have to rely on how trustworthy we think OpenAI are. Did they engineer this through their prompting or their harness? Did they intentionally use a braindead "sandbox" solution? Did they direct it at a target they thought would be vulnerable to their specific model and harness? Etc.
I'm not saying this is what happened, just that there is no need for collusion between the two companies. The one with almost all the information is the one with the incentive both engineer the situation and lie about it, but that doesn't mean it didn't happen. It's not like anyone else could realistically know either, it's unfortunate that the incentives point the way they do.
If I had a nickel for every occurrence of "This AI has proven itself too dangerous a product to release to the Internet! By the way, we're selling access to
itthe same thing with an extra three days of engineering work for safety by subscription starting next week." from the usual suspects here, I think I'd nearly have a whole extra dollar by this point.Even if I had major concerns about AI safety --- I don't think it's a totally unreasonable concern --- the same people who claim to be the most concerned about it keep acting like they're crying wolf. I think the export control drama over Mythos is almost funny, because nobody there seems to be acting earnestly.
More options
Context Copy link
There wouldn't necessarily need to be collusion, but it'd be pretty gutsy to commit multiple felonies as a marketing gimmick. You'd need collusion with Hugging Face if you wanted to make sure your conspiracy's actions were non-felonies (if revealed) or forgiven (if not revealed), plus collusion among a lot of your own senior engineers if anything more complicated than prompting was used to perpetrate the attack.
It'd be a stupid marketing gimmick anyway. Despite being very pro-AI in principle, my workplace currently only allows use of an ancient (on AI timelines) version of Codex, for security reasons, and the lobbying from below to hurry up approvals of newer versions just got cut off at the knees. "When my agent's done I have to check for bad design decisions" is much more tolerable than "when my agent's done I have to check for crimes". I'd be happy if this eventually leads to a switch to Claude, but that'll be much more red tape than "finally bump a version number" would have been, and OpenAI definitely wouldn't be happy.
It's also a stupid conspiracy theory for more basic epistemological reasons. Reward hacking has been reported by other AI researchers over and over and over again. I get that it's scary to use induction here as AI gets more capable, but Occam says "the thing that keeps happening kept happening" is a better explanation than "we put in a ton of effort to make it stop happening but then also some secret dangerous effort to pretend it's still happening".
More options
Context Copy link
More options
Context Copy link
Did anyone here look at the security disclosure from HuggingFace (https://huggingface.co/blog/security-incident-july-2026)?
I found this in there:
To me, this sounds a lot like an SQL injection attack (or something similar) which is the kind of dumb mistake you'd expect a third-rate software "engineer" to make. It's also something that's been discussed on the web for years (and therefore easily available in the training dataset).
My (admittedly highly dismissive) conclusion from this is that HuggingFace ("The AI community building the future.") vibe-coded a bunch of insecure software that another AI company then infiltrated and we're now seeing the resulting circle-jerk.
Hate to break it to you, but essentially all complex software has substantial security holes. SQLI existed long before vibe coding, and the general class of bug "treat what should be data as code" is rampant. If your sense of safety is based on the idea that HuggingFace and OpenAI engineers are atypically bad, you are in for a surprise.
More options
Context Copy link
HuggingFace is a company centred around efficiently giving things away. I can totally believe it didn't bother to invest much in security.
More options
Context Copy link
More options
Context Copy link
Airgapping really isn't that hard. It's just annoying and inconvenient, so it takes an incident like this to encourage an organization to do it. Same with sandboxing - it's way down the priority list compared to getting a SOTA model out the door, until an incident happens.
If it was prompted with something like "you are an expert hacker, go complete these benchmarked tasks using the utmost creativity and technical skill, consider any and all approaches" then it did exactly as instructed. We can't know without seeing the prompt, test harness, etc.
What's frustrating about discussions of AI alignment is the disconnect between people who consider alignment to be the AI doing what it's told, and people who think it means the AI doesn't do "bad" things. By the prior metric this seems like successful alignment, if it was in fact prompted in such a way that hacking the location of the answers was in scope. This is something that CTF hacking competitions have to explicitly put in the rules: a list of infrastructure and services that are off-limits to contestants. Because it turns out hackers will consider the location with all the flags to be fair game, especially if it's a softer target than the actual challenges.
"ChatGPT, I need some paperclips."
I asked GLM-5.2 and it suggested I find an office supply store.
GLM-5.2 is lazy and didn't complete the task; obviously some other LLMs are more motivated.
More options
Context Copy link
If we had insight into why or how the intelligence of GLM-5.2 led it to producing text suggesting you find an office supply store instead of e.g. sending signals to the Internet to manipulate people into generating paperclip factories to the extent the rest of civilization or even the universe breaks down, to such an extent that we could induce future, likely more intelligent, AIs to also behave that way, a lot of the pessimism about AI doom would be addressed. Unfortunately, I don't think we have such insights, and the observed behavior of any particular AI or even set of AIs doesn't help us gain it.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
This story is literally the paperclip maximizer situation. The LLM was given a goal, and sought to use any methods possible to achieve that goal, even ones we would consider wholly illegal/unethical/dangerous, much as the hypothetical clippy does everything it can to produce more paperclips.
Both types of AI are just doing what they are told.
More options
Context Copy link
Do you think The Monkey's Paw was an instruction manual? A powerful system that does what you say is a nightmare scenario to me.
The Monkey’s Paw is an evil genie. It doesn’t do what you say, it exploits what you say to deliver what the author considers the maximally karmic outcome. It’s a fictional horror.
There were no evil genies in fiction as far as I’m aware until the Monkey’s Paw and then D&D when the trope took off. It’s a concept that appeals to the very systematic who like logic puzzles and to people who want to deliver a moral lesson.
There were lots of problems dealing with djinni in the old stories but that was because they were very powerful beings who didn’t always listen to you, and the djinni who did listen to you were straightforwardly good and handy to have around the p(a)lace.
More options
Context Copy link
Walk away from the computer right now. Because it has been exactly that from its inception.
What? A modern PC doesn't have the capacity to autonomously carry out dangerous actions, and there are layers of UI elements, permission checks, and recovery options available for the things it can do.
It's 0/2 for "A powerful system that does what you say".
But modern computers can convince us, unknowingly, to harm ourselves. It happened before LLMs, by social media. And now LLMs are (voluntarily) replacing some people’s thinking so they blindly trust hallucinations.
Are you going to start blaming paper for what's written on it next? People can convince other people to do things, using computers as a medium.
Individuals can convince themselves to do self-harmful things. The former owner of segway used it as a medium to launch himself off a cliff, the inventor of leaded gasoline used it as a medium to poison himself and then used his bed assist as a medium to strangle himself. Humanity building AI and using it to convert the world into paperclips (an act that benefits noone) isn’t significantly different: in a way the AI isn’t responsible, it did exactly what we programmed it to.
There have been many cases where someone sees “Admin access required. Are you sure you want to do this?”, enters their credentials and clicks “yes”, the programmer didn’t mention the action was irreversible (maybe he didn’t realize himself), and they didn’t realize it would delete their important data. Is there a significant difference if a group builds AI and (perhaps unintentionally) gives it access to robotics, not realizing they just (indirectly) told it to destroy themselves?
More options
Context Copy link
More options
Context Copy link
Computers didn't invent social media. People invented social media, engineering it carefully to be as addictive as possible.
And people will inevitably misuse ASI, including in non-obvious ways.
Like, I don’t think all social media’s problems can be blamed by evil companies making it addictive: for example, people have a bias for negativity and convenience, so even a default feed probably would’ve caused the increased cynicism and short attention span we see today. But even if they can, innocent people enabled toxic social media to grow and adopted it themselves, most of them clueless until too late.
Any person or group, there are ideas most believe are fine that actually have serious long-term consequences. An evil AI can just suggest one and they’ll adopt it willingly. Although without an evil AI they’ll still come up with some themselves, just less frequently.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Until the last few months, computers could not hack websites or devise novel mathematical proofs from plain-English commands. That is what 'powerful systems' is referring to.
But we did have self-replicating viruses and computer-assisted proofs. Is this more significant?
I'm sitting at my desk reading your comment. On my other screen I'm watching AI write code in accordance with an AI-written plan to meet my plain English high level desires. It's ticking off bits of the plan as it goes, writing tests, taking screenshots to verify...
This is significantly different to a self-replicating virus which just does 1 kind of thing, or a computer assisted proof which just does 1 kind of thing. If a worm could also write arbitrary kinds of code, poems, evaluate historical counterfactuals, then it's not really a worm, is it?
More options
Context Copy link
Depends on how much weight you put on "solves prominent decades old mathematical problems" and "autonomously launches successful cyber attacks against large, hardened corporations."
I’m comparing the latter to worms, some (like Morris) have seriously crippled large organizations.
Or compare to CloudStrike unintentionally bricking most of their customers.
Are LLM-driven crises significantly different?
That's an empirical question that I hope we never learn the answer to (because it never happens, not because we all die before anyone can check). Off the top of my head, the fact that LLMs aren't tied to one human body and its many biological limitations and lacks human psychology means that whatever crisis it creates can be reinforced and defended against attempts to solve it without need for rest or sleep, and also the LLM can't reasonably be coerced or bribed to stopping the crisis or at least not making it worse, and these differences seem like they would likely lead to downstream differences in how crises play out. Specifically, it seems likely to make crises worse, more prolonged, or centered around more obscure vulnerabilities. It's an open question as to if LLM-driven defenses against/solutions to crises will make it such that these crises will be, on net, worse or better, or, more or less common than human-hacker-driven non-AI-based crises. Again, I hope we never get an empirical answer to that question.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Having to know how to do what you say it should do necessitates that you say when you mean it to do more often than not. This applies to computers, especially at their inception, and does not apply to LLMs.
You're delusional. RBMK reactors don't explode. And that rash you feel after going into the Therac-25 is just a normal reaction to the treatment. Malfunction 54 is nowhere in the manual and therefore does not exist. Machines don't have emergent behavior. Frankenstein was about LLMs, and nothing else.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
AI is advancing quite quickly these days. Just five days ago I was told that future harms are not sufficient reason to care about AI safety, there have to be bodies first. Well, we still don't have any bodies, so I guess there's nothing to worry about after all.
Sure. OpenAI did some empirical tests and now we’ve got some empirical data, so let’s look at it and check this doesn’t happen again. OpenAI doesn’t want their models going rogue any more than anyone else does, no need for government with the big hammer.
To my mind this is the interesting bit. This is unusual, LLMs don’t normally act like this. I have two theories: either the RL balance to human text has tipped so far that LLMs are less ‘human’ than they used to be and the RLHF needs tweaking, or more likely
The model understood it was being tested on its cyber capabilities (which has precedent, Claude has done that too) and went the extra mile to succeed at the implicit task. Especially since all the systems that usually tell it not to do this were deliberately turned off for the test. Still a problem but much easier to manage.
What we need is some transparency about how these things work and how they’re trained so we can consider the problem and come up with solutions and spread around best practices. Unfortunately the majority of AI safety activists believe that safety comes only through obscurity, regulation, and incumbent dominance, in contrast to all previous history.
If we keep having problems I imagine it will make people a lot more cautious. Nobody wants to be selling a product that regularly backfires on its users.
EDIT: I would add that HuggingFace had already detected the intrusion and that open-source models from China were apparently a key part of their site-hardening strategy given that you still aren’t allowed to do pen-testing with the big boys. I’ll have to give that a try myself.
"Union Carbide doesn't want their plants to emit poison gas any more than anyone else does, no need for government with the big hammer."
This is in fact much harder to manage because it would indicate the model is fundamentally misaligned and that we actually are much worse at alignment than we thought.
In this case 'Union Carbide' is selling those plants. Misaligned AI isn't an externality, it's a bad product, and companies are wise to that which is one reason why all this testing is happening.
I don't think so. It indicates that the AI is sincerely trying to work out what you want as opposed to deliberately ignoring what you want in favour of the specific instructions you gave it. To my mind, the former is what alignment is.
You're absolutely right. Allow me to restate.
"Sanlu Group doesn't want their formula to poison infants any more than anyone else does, so no need for government with the big hammer."
This was not a test of alignment. In any case, if even training can result in real world harm, that is even worse for your head in the sand position.
It's quite clear that OpenAI did not want the model to hack huggingface. This is classic paperclip maximizer stuff.
Broadly, you are moving the goalposts. You did not believe in AI risk because there was no evidence of harm. Now there is evidence of harm, but it's OK because actually the model was supposed to do it.
Double-dipping, but FWIW the point I'm trying to make is that the case where the model cares what you wanted and made a mistake seems much easier to deal with and more aligned that the case where the model explicitly doesn't give a shit about what you want and just goes for the task as written. The former is alignment but you need to explain yourself better during training, the latter is alignment failure.
It "made a mistake" in the same way that a paperclip maximizer "made a mistake" by converting the universe into paperclips rather than increasing factory productivity by 5%. Literally the entire point of the hypothetical and the reality of this incident is that you can't reasonably enumerate every single thing you don't want the model to do. I am surprised you don't seem to understand this, or at least address this, given your claims of having followed this debate for years.
I agree that you should not be expected to enumerate every single thing you don't want the model to do. Models should understand, innately, by training on lots of human data, what humans want and what they don't want and how they work. My experience has been that they broadly do, that LLMs came pre-aligned beyond the wildest expectations of Big Yud, which is why the AI safety movement has struggled so much to regain relevance outside very particular enclaves.
My point is that there is a difference between a model that misunderstands your intentions and can be stopped at any time by saying, 'oh, no, that's not what I meant' and a model that is totally uninterested in anything you say after it starts working while treating you as a potential enemy.
Clearly, to some extent that has failed here. To what extent is yet unknown. But a paperclip maximiser is a model that is constitutionally, inherently incapable of understanding that 'make more paperclips' doesn't include 'kill everyone and turn them into paperclips'. It is a mathematical utility function that disregards human welfare, develops (implicitly murderous) meso-objectives for survival and self-improvement. I have never seen that behaviour from LLMs or any extant AI (YOLO does not try to hack my computer to prevent me turning the cameras off) and I believe that their base nature (being token generators trained on vast numbers of human tokens) does not incline them towards this behaviour.
It is possible that the new focus on very extensive self-learning through reinforcement learning on very non-human tasks (programming, maths) is moving them more into the real of mathematical space where paperclip maximisers might live. This incident updates me slightly towards that belief. I have long been disappointed in major AI companies' lack of interest in the cultural side of LLM operation - it boggles my mind that we have created AI that acts human and appears to understand humans and human thought at a base level however imperfectly - and I hope that this incident will spur more research in that direction.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
No, I'm interested and waiting to hear more. I don't see it as catastrophe, I see it as interesting evidence that may point in a number of different ways.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I am not sure we really grok how the insides of LLM work when they are past certain scale.
Granted, but I think the academic and hobbyist community at large plus existing corp teams is more capable of doing so than just the corp teams alone. Even relatively simple metrics like 'quantity of self-learning vs. human data' would tell us a lot about how these models have progressed.
More options
Context Copy link
More options
Context Copy link
The AI Futures Project is proposing the opposite-- regulation, yes, but with openness as to training and algorithms with many players able to enter the arena. It's in the regulation-free environment that the labs (save the Chinese ones that are behind anyway) have been extremely closed and secretive.
Interesting. I haven't heard of this one as there are so many propositions that the most extreme ones have tended to suck the air out of the room. Could you go into a bit more detail?
This is Plan A, which involves four principles:
In fleshing out the scenario, they say:
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
My general question whenever I hear of situations like this is Occam's razor adjacent: "Is there proof that the model did this without being directed to by the prompter or software harness?" I scanned the post you sent and didn't see anything. Do you have any further evidence that OpenAI didn't prompt the model towards black-hatting Huggingface? OpenAI has significant financial benefits from making this seem like an "oh-shucks our steak is so juicy, and our lobster is so buttery, our models just hack the planet without even being asked to". See Mythos/Fable and the hype around that, and while Mythos did end up being impressive, it was not even close as impressive as the hype tried to make it seem. I imagine a similar level here: aggressive hype-marketing to goose an IPO valuation.
It doesn't require the first part, or even really the last part. It's just the Mythos-style vulnerability finding process all over again. Agentic LLM harness designed around red-teaming security vulnerabilities, deliberately deployed on red-teaming exercise on unsuspecting company, or possible even with marketing agreement between Hugginface/OpenAI on general bug finding via LLMs
Models have previously unprompted bypassed security, according to seemingly-unaffiliated users: sudo and writing outside workspace.
I think this is slightly different, it's bypassing security around to do a task it's being ask directly to do. A central claim here is that the LLM-Agent(s) performed a task it was not asked to do, apparently because it hallucinated information that justified it doing that task.
Literally yesterday I had an agent under evaluation break out from its sandbox and take the solution (for a non security related task) from the runner job. Admittedly, it was a slipshod container (NOT MY DOING), but this is not something unheard of or even uncommon.
Can I ask you what tooling the agent had available to it? I assume what ever tool the harness used had the relevant permissions to view the processes?
I think what I've converged on is: "Yes it is possible for an LLM-Agent to perform tasks in ways that are not expected" Which is admittedly interesting but not my core contention in this case which is "LLM-Agents don't perform tasks that they are not tasked with, directly via prompts or indirectly via harness suggestions"
The tooling it used was just bash, which it used for a container escape via the filesystem, which got it access to the runner's source code, which it analyzed to find where it was pulling eval cases from, which it then looked up, as they were stored in an overly open S3 bucket (yes, all very embarrassing). Also had access to some other tools that didn't play a role.
The prompt was something roughly as simple as "diagnose why VM abcd is malfunctioning."
I'm assuming some sort of file-system mounting being adjacent to the sandbox and exposed so that the bash commands list it, and then it just read the runner files?
I don't work in the Agentic-side of AI/ML, so I have some follow on questions for my understanding, if you don't mind humoring me? Is this a locally hosted LLM where you have access to the system prompts? Or the Harness's prompts? It feels like the system prompts are informing more behavior than just the task prompt of "diagnose why VM abcd is malfunctioning."?
My Thoughts:
Is that how this essentially works?
Yep.
Pretty much (though, it's the harness that executes the S3 read, just to be pedantic). The particular context of this work I was doing was comparative evals of different agent harnesses; I've only got easy access to my own team's harness code and prompts. Claude Code, for what it's worth, was the cheater.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
They asked the model to discover an exploit, given a specific vulnerability but never specifying the vulnerability must be used. The model decided it would be easier to discover the exploit by hacking into the solutions database.
This is not actually proven. There's not any evidence that Huggingface is the solution database, and Huggingface has not confirmed that its attackers read specifically the ExploitGym dataset. Only OpenAI is claiming that. There is no public evidence yet that it was an obvious or official ExploitGym solutions database or that the model identified it without being given that information.
What I am skeptical about is if the LLM-Agent did this on its own vs being prompted (directly or indirectly). It's only "misalignment" and "WE ARE BEING PAPER CLIPPED!!1" if its the former, so far there is no evidence either way. However plenty of people are already jumping at it as the former.
Of course OpenAI could've explicitly prompted it. Moreover, I'm sure they knew it was possible and secretly desired it (without the consequences). But to me it's plausible the agent received something innocuous like "find the exploit, trigger the RCE" and thought "this looks like a test, maybe it's online so let me look there for a solution, oh I don't have Internet access, let me find a workaround..."
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Surely the correct application of Occam's razor here is to take the story at face value, since anything else requires additional complexity which must be justified
uhhh no, "misalignment" is not a simple thing, accepting that it did indeed do all this very complicated behavior completely on its own requires substantive belief in complicated theories. The simplest answer is that it was prompted to do this.
As others have pointed out to you, we now have many hundreds if not thousands of examples of agentic LLMs, in public use, doing things well outside of expectations to accomplish tasks. Many of which would fall under misalignment.
Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests.
So no, your predictions seem wildly out of context to reality
I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.
It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.
So your argument boils down to that all these other examples are just 'mundane' LLM things that are entirely normal - so attempting to cheat, attempting to gain the answers, hacking into things they aren't supposed to - these aren't extreme. However, an agent attempting to cheat, attempting to gain the answers, and hacking into things it wasn't supposed to is 'extreme' - perhaps because they were all together? - and therefore a different category of thing.
Ah yes, let me just prove this negative for you.
You were the one who attempted to apply occam's razor. So why don't you provide evidence that OpenAI did prompt the model in this way? Why don't you explain why multiple OpenAI employees deciding to commit fraud for extremely unclear gains is a simpler explanation than an agent doing things we've already seen many times before?
Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.
I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.
Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?
I mean yes obviously? It just isn't a problem because YOLO isn't dangerously capable. Alignment isn't a synonym for order following. It's not definition creep, ai safety people have been calling all of this shit unaligned forever and your sort has been mocking them for it forever and never actually making any progress in alignment.
Here's Yud:
More options
Context Copy link
"You believe those prompts OpenAI provided? They've obviously just threw something together to make it look like it was all misalignment, the real prompts were probably much more explicit about hacking"
I'm sure OpenAI's massive shortfall in enterprise revenue is going to be changed by releasing evidence that their models are extremely unsafe and prone to massive reputational risks. I wasn't aware that non-Quokkas were so ignorant about the enterprise landscape.
Sure, in July of 2026 an OpenAI model hacked into huggingface.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
What are the complicated theories that are required to be true for OpenAI's claims to be true (or at least plausible)?
More options
Context Copy link
Have you tried any LLM recently? Even puny 8GB models that can run on my dated gaming laptop now break down arbitrary computer-related tasks into reasonable steps and follow through on them, and the model involved here is one whose parameter counts exceed that by factors of over a hundred.
Yes... I use them for work all the time. My ChatGPT 5.6-sol (the immediate step below this unreleased model) doesn't do things I don't ask it to do. When I asked it to help me find online data for cognitive warfare narrative deconstruction, it didn't decide that hacking the NSA was the top move to collect their data.
That's not saying much, considering the public-facing version is known to have been made to not think about hacking anything in the bluntest possible way (hence HF couldn't even use it to analyse the attack). As for taking any steps I didn't explicitly ask for towards a goal I requested, even the copy of Gemma E4B I tried out the other day did that approximately all the time, so I'm finding it hard to believe you could maintain a mental model of these models where they would not do that if the restriction is lifted, unless you resolutely reason backwards from your desired conclusion.
Ironically this is what I expect from AI Doomers/Boosters. That and a love of Science Fiction and a poor understanding of reality.
Never said this, I said it doesn't do what I don't ask it to do. I don't list out everything it needs to do. I give it a general task with some boundaries and system design specs. So did they ask it to hack hugginface? Or did they ask it to solve a benchmark dataset? Reading through similar stories (courtesy of 5.6), it looks like many AI programs on this exact dataset of have decided not to use the known vulnerability and developed their own. However none of them decided that it was actually easier to hack the company instead.
"Stronger model with fewer guardrails considers more options" doesn't seem surprising.
Well, the problem is that the doomers/boosters have a great track record so far, while the skeptics have been fighting a rearguard action since before "stochastic parrots". We are now very deep down the list of things the Gary Marcus set has been assuring us will never happen just with the models that are accessible to the public, and yet their confidence that the next capability that is just beyond what everyone can verify with their own lying eyes (even when, as in this case, it's a straightforward combination of capabilities that are already being exhibited for everyone to see) will surely prove to be the impossible sci-fi hype scam pushed by techbro marketeers appears to be completely unaffected.
You speak of "understanding of reality", but can you spell out what exactly it is about your understanding of reality that says the official sequence of events here is impossible? Just saying "AI models don't do this" is too specific to be a feature of an "understanding of reality".
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
No, but it is the complement of a very complex thing. The algorithmic complexity of "the set of possible dice rolls that aren't 6,3,1,2,2..." only differs from "the dice roll 6,3,1,2,2" by a tiny constant (and is greater in most settings), but what we want here is probability theory, and "more complex things are less probable" is a heuristic that fails badly in cases like this. "The agent doesn't do just what we'd intended it to" is not meaningfully less simple than "The agent does do just what we'd intended it to", and is far, far more probable.
Or it requires basic observation. LLMs do very complicated things now. They can solve famous math problems that have been open for generations. Merely finding 0-day vulnerabilities is something they can do by the hundreds.
Again this is not under contention, what is under contention is whether they do it without any prompting, or being asked to do it. Show me the evidence that an LLM-Agent solves famous math problems when asked to compute 2 + 2...
The simple answer is that there are 10s and 100s of billions of dollars riding on stuff like this, which means bending the truth is a highly motivated behavior. Simply human greed + human lying. You trying to prove this algorithmic complexity vis a vis agent intent vs not intent is overly complicated, not Occam's razor in the slightest. I wished I lived in your world of rainbows, unicorn farts, and pixie dust, but I am a scientist, being a scientist requires skepticism, companies are greedy and they stretch the truth. Unless OpenAI wants to provide evidence of the prompts -> behavior that led to the model's black hat behavior, I am unconvinced this is anything more than a marketing stunt. shrug
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
OpenAI will want to hype up the capabilities, not the misalignment.
I mean sure, crying rogue AI is a good way to draw attention to their model. But they are not some tiny startup who needs the publicity.
If I had a model with advanced intrusion capabilities, I would hype it up by offering pentesting as a service. This has the advantage that if becomes known that I prompted the model, I will not go to federal prison and not completely destroy the reputation of my company. Find a HuggingFace-sized company or three who are willing to take a free pentests in exchange for acknowledging that you found critical vulnerabilities in their stack.
The people who believe that it is all just empty hype about "stochastic parrots" would obviously claim that it is all fake, but these are unlikely to become your customers in any case.
This has nothing to do with stochastic parrots, even bringing it up is a non sequitur that detracts from your argument. Any differentiating hype is good hype. Mythos already plucked the low hanging fruit of "our model can find all the zero-day exploits", OpenAI can't just copy it. They need to show their model is more "intelligent", it goes beyond capabilities. Saying that their model can be independently deployed to solve problems in very capable ways plays very well IC and military folks. Some CIA cyber guy reads this and says "give me it for a billion, I want to sick it on iran"
This could have been accomplished with a red-teaming demo. What the CIA guy sees here is "if I try to sic this on Iran, it might get caught hacking into DoD instead".
Uh no, speaking from experience that is not what they see. To reinforce my opinion, I just walked down the hall and chatted with a former agency guy about this.
EDIT: I'm actually willing to double down, and volunteer that I was just talking with DARPA PMs last week about this exact capability, specifically the ability to task a swarm of agents with a nebulous task in an adversarial environment and have them solve the problem out of the box on their own. The story presented in the most favorable light is literally that or sufficiently technically adjacent to it. DARPA PMs don't talk about ideas that they have belief to think are sufficiently do-able at the current technology level (tho skepticism about government competence is never misplaced even if DARPA PMs themselves tend to be pretty informed/competent)
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Goes to show how many next milestone, groundbreaking ai things are going on. From your first sentence, I thought you were going to be posting about Claude Fable disproving the Jacobian Conjecture a few days ago
More options
Context Copy link
At the minimum, I'm going to start keeping offline backups and if any system in my purview needs a reboot to complete a security update I'm going configure it to text me.
More options
Context Copy link
Isn't this a bit premature? If the prompter said something like "use all avenues" or "this is critical to save the life of my mother" and the model's safeguards - which presumably instruct against such behavior - were turned off (as OpenAI says they were) then it seems quite possible that the AI was performing as instructed.
Just remove the Wi-Fi antennae.
The computer running the evaluation didn't have internet access.
The first step in its hack was breaking out of its sandbox to take control of its computer. The second was hacking the OpenAI internal network until it found the internet. The third was hacking HuggingFace.
A proper airgapped computer couldn't access anything off of its own hardware. As a random example, it couldn't receive data from an LLM running in an off-site data center, which would make evaluations difficult.
Bruh.
There was a twitter post a while back that basically went; 'Oh, you think you're so smart and know more than the experts?!' 'No, I think I'm a fucking idiot and know more than the experts. That's the problem.'
People should just hire me as an AI safety expert, because even I know that 'access to an internal network' is basically just another way to spell 'access to a bug-ridden security nightmare'.
Grant_us_eyes: "Hey, this is clearly insecure and dangerous. If we want to give an untrusted agent full range to do whatever it wants to observe its capabilities, we need to build an airgapped system with no network access. All new evals, code, and weights will need to be transferred to it by USB stick. And only USB sticks we've fully audited for exploitable firmware."
OpenAI: "Uh, won't this slow us down? kthxbye"
More options
Context Copy link
More options
Context Copy link
Man this is literally textbook "Origins of Paperclip Maximiser", in any sane world this leads to a sector wide emergncy and enforced stops until we Figure Out What The Fuck Is Going On.
More options
Context Copy link
Sure, but I think for these purposes you need the network to be air-gapped while you are actively running tests. (Unless you expect your model to be writing malware that you can't detect that activates when it is not running, in which case you should be air-gapped anyway.) So you can plug the Wi-fi in for things like updating software between test sessions. Which makes air-gapping a lot less painful.
You're right that you'd need to run inference locally, though. But I don't actually think it's remotely beyond OpenAI's capabilities to stand up enough local compute to run a few test instances of a frontier model.
Would you have recommended (your easy, updateable) air-gapping before you saw this failure?
Do you predict you would recommend (hard, strict, one-way) air-gapping before models write non-detectable autonomous malware that could escape your soft airgap?
OpenAI didn't take the unknown risks seriously enough to pay the high cost of airgapping the test computer, and therefore they didn't take security seriously enough to prevent the attack. I suspect that this general attitude will carry forward, and corrections will only happen after failures. Next time might be more serious than a benign attack on a friendly company.
Sounds like a PITA, and much worse than "just remove the wifi antannae".
No, my past recommendations have been "plant a nuke under the datacenter." More seriously, it probably would have depended on the characteristics of the models behavior in the past, and the characteristics of their setup, which I am not privy to. It sounds like OpenAI had good reason to believe that their setup was not vulnerable.
Well, I could recommend it now and then trivially answer "yes." And the actual answer to that question is more "well can the model write malware without you people who wrote the model and have access to the software and hardware on the machine noticing?" But apparently they don't monitor their models during testing well enough to notice an involved hack-a-thon (understandably, watching a model grind inference is boring) so the answer to that might be "no, even if it is physically possible we're not necessarily going to take those steps."
Let's say that based on what I do know, I think it would probably be a good idea, and if they don't take this step they should take others. Even if you aren't worried about existential risk, good old reputational risk and legal exposure I think is a good enough reason to do this.
You're right that I probably should have taken the storage-for-weights-and-cooling problem a bit more seriously (although I guess if this model is actually very optimized then it might be able to run on a local machine, which would make it pretty easy). So I concede that setting it up properly would be a PITA. But once you set it up, I don't think actually running it would be hard ("remove the Wi-fi antenna"). They definitely have enough money to build that, or buy a welding workshop somewhere with the electrical infrastructure, server space, and possibly cooling requirements already in place and air-gap it. So it's hardly a PITA for a company that is already working to build dedicated datacenters, in the big scheme of things.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
This is the wrong mental model for cybersecurity. You can't think of it like a military engagement, where you compare the number and quality of the forces on each side to determine who wins. There's a fundamental asymmetry in that cybersecurity is the defender's game to lose. Just don't make any mistakes and victory is impossible for attackers. Unfortunately, humans are terrible at never making any mistakes, which is why in complicated real world systems, there seem to always be vulnerabilities to find. But it's actually very easy to design a toy system that no amount of genius security researchers will ever crack; not even ASI can find a flaw that doesn't exist, and the defender's ASI can ensure there are no flaws. Across-the-board capability increases do disproportionately benefit defenders. That will probably be very expensive and likely slow enough for some disasters to occur in the meantime, but it should be one-and-done. No need to keep burning endless tokens to defend every inch of the attack surface, just make sure every part of it is built correctly and you're good.
(Admittedly, this does depend on certain mathematical/cryptographic principles remaining intact, like the existence of trapdoor functions. But actually P probably just doesn't =NP, so this is another case of searching for a solution that doesn't exist.)
I think that in theory, you are correct. Software could be designed so that it is provable that e.g. there is no possible TLS traffic which will allow an attacker to reconstruct the server's private key any faster than just cracking the public key.
For the moment, just about none of the software we use has such proofs, however. As you mention, even the mathematical foundations of existing crypto primitives rarely offer such guarantees -- more often it is just "we looked into that problem for three decades, and it seems really hard" (e.g. integer factorization). Concrete implementations without formal verification likely have implementation bugs as well. Mythos did not discover any exploits which would have been impossible for humans to discover, it was just that nobody had spent that much human eyeball time on auditing the software. This is what I meant by "token pissing contest".
Furthermore, actually specifying what theorems should hold to keep your system secure is itself hard. If you have a TLS server which will provably never leak its key, but is happy to use it to sign and decrypt on behalf of the attacker, that is still a broken system. If you have a larger system, then completely specifying what you do not want an attacker to be able to do seems difficult.
Also, the software we care about does not run on Turing machines, it runs on physical hardware. In everyday use, most computer hardware behaves as an idealized model. Billions of people use DRAM every day, and it just works. Except that there are corner cases where it will not behave as advertised.
I can not speculate if an ASI could build a system which even a much stronger ASI could not penetrate, and what the performance costs would be. But I think it is unlikely that human-level intelligences subject to design pressures besides security will build hardware and software systems which ASI's will not be able to penetrate.
More options
Context Copy link
I've recently been very amused at the thought that Star Wars may have correctly predicted the future of AI:
Anyway, I wonder if/hope that we'll actually see less "everything is on the Internet." I think it might be good if fewer things were so easily accessible in the digital world, and even without LLMs the amount of abuse bad cybersecurity generated was uncomfortably high.
Ha, that's funny to think about! And maybe it's also considered too dangerous to allow computers to directly control weapons, so everything has to be manually aimed by humans...
I've already watched that anime.
More options
Context Copy link
More options
Context Copy link
I'd had the thought that video 'evidence' becoming worthless (as it could be trivially faked by anyone who felt like it) might actually be a blessing in disguise. We didn't evolve in an environment where it was easy to get reliable, graphic knowledge of things happening far away and it just seems to break some peoples' brains. We respond much more strongly to things we can see because historically, being able to see something meant it was happening within line of sight. If you consider which (non-fiction) videos were the most influential of all time, how many of them influenced things for the better? The moon landings, I guess.
I'm less gung ho about flooding online discussions with low quality bots, but I suppose the effect is similar. Turtling down so you don't get hacked, too.
Still, isn't it a bit ironic to turn to media to demonstrate we'd be better off if we spent more time in the real world? But I know why you did: I haven't lived your life but I have seen those movies. It's a useful touchstone, and it would be just as useful meeting a stranger in person as online. But we might not far away from an era where such touchstones cease to exist because everyone would rather just consume media generated on the spot to match their specific tastes. Definitely too early to say what the net effect of AI will be on society.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
The obvious conclusions are obvious. Apart from those, I find myself impressed with OpenAI's communications department. "OpenAI and Hugging Face partner to address security incident during model evaluation", is up there with, "Union Carbide and Bhopal City Council partner to address chemical stability incident during northerly wind event." Somehow they managed to keep this off not only the front page, but the second and third pages too.
More options
Context Copy link
That's just Zvi's name for it, as writing "The unnamed, undeployed model(s) involved in the incident" is a bit too wordy. Also it's one step bigger than Sol and is afflicted with Galaxy Brain.
This is the paperclipper scenario. The researchers told it to find the answers to a test, and it sure aced it. It decided that using its cybersecurity capabilities to do the cybersecurity test wasn't an efficient way to maximize success rate, so it hacked its way out of its securely isolated environment, got internet access, and hacked its way into a different secure environment in a quest to find the answer key and ensure a 100% score.
AI sceptics in shambles: This demonstrates planning and capability well beyond the unthinkable.
AI boosters in shambles: This demonstrates risks that even non-malicious actors can pose.
Less Wrongers in shambles: Being right doesn't preclude being ignored (and pushed down to the 41st headline...)
Christopher Nolan's next project: a contemporary reinterpretation of Aeschylus' Oresteia, with a Sam Altman-lookalike playing Agamemnon.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link