So your argument boils down to that all these other examples are just 'mundane' LLM things that are entirely normal - so attempting to cheat, attempting to gain the answers, hacking into things they aren't supposed to - these aren't extreme. However, an agent attempting to cheat, attempting to gain the answers, and hacking into things it wasn't supposed to is 'extreme' - perhaps because they were all together? - and therefore a different category of thing.
Provide me actual evidence that OpenAI did not prompt the model to act the way it did.
Ah yes, let me just prove this negative for you.
You were the one who attempted to apply occam's razor. So why don't you provide evidence that OpenAI did prompt the model in this way? Why don't you explain why multiple OpenAI employees deciding to commit fraud for extremely unclear gains is a simpler explanation than an agent doing things we've already seen many times before?
As others have pointed out to you, we now have many hundreds if not thousands of examples of agentic LLMs, in public use, doing things well outside of expectations to accomplish tasks. Many of which would fall under misalignment.
Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests.
So no, your predictions seem wildly out of context to reality
What's frustrating about discussions of AI alignment is the disconnect between people who consider alignment to be the AI doing what it's told, and people who think it means the AI doesn't do "bad" things.
This story is literally the paperclip maximizer situation. The LLM was given a goal, and sought to use any methods possible to achieve that goal, even ones we would consider wholly illegal/unethical/dangerous, much as the hypothetical clippy does everything it can to produce more paperclips.
Both types of AI are just doing what they are told.
Surely the correct application of Occam's razor here is to take the story at face value, since anything else requires additional complexity which must be justified
A number of the most prominent AI voices on xitter are OpenAI people speaking personally; it very much appears that OpenAI doesn't really interfere at all as to what their top people are allowed to speak about on social media. And Ball is a very recent hire for them and has a long history of posting takes just like this, so there's nothing out of the ordinary there.
- Prev
- Next

"You believe those prompts OpenAI provided? They've obviously just threw something together to make it look like it was all misalignment, the real prompts were probably much more explicit about hacking"
I'm sure OpenAI's massive shortfall in enterprise revenue is going to be changed by releasing evidence that their models are extremely unsafe and prone to massive reputational risks. I wasn't aware that non-Quokkas were so ignorant about the enterprise landscape.
Sure, in July of 2026 an OpenAI model hacked into huggingface.
More options
Context Copy link