This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Surely the correct application of Occam's razor here is to take the story at face value, since anything else requires additional complexity which must be justified
uhhh no, "misalignment" is not a simple thing, accepting that it did indeed do all this very complicated behavior completely on its own requires substantive belief in complicated theories. The simplest answer is that it was prompted to do this.
As others have pointed out to you, we now have many hundreds if not thousands of examples of agentic LLMs, in public use, doing things well outside of expectations to accomplish tasks. Many of which would fall under misalignment.
Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests.
So no, your predictions seem wildly out of context to reality
I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.
It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.
So your argument boils down to that all these other examples are just 'mundane' LLM things that are entirely normal - so attempting to cheat, attempting to gain the answers, hacking into things they aren't supposed to - these aren't extreme. However, an agent attempting to cheat, attempting to gain the answers, and hacking into things it wasn't supposed to is 'extreme' - perhaps because they were all together? - and therefore a different category of thing.
Ah yes, let me just prove this negative for you.
You were the one who attempted to apply occam's razor. So why don't you provide evidence that OpenAI did prompt the model in this way? Why don't you explain why multiple OpenAI employees deciding to commit fraud for extremely unclear gains is a simpler explanation than an agent doing things we've already seen many times before?
Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.
I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.
Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?
"You believe those prompts OpenAI provided? They've obviously just threw something together to make it look like it was all misalignment, the real prompts were probably much more explicit about hacking"
I'm sure OpenAI's massive shortfall in enterprise revenue is going to be changed by releasing evidence that their models are extremely unsafe and prone to massive reputational risks. I wasn't aware that non-Quokkas were so ignorant about the enterprise landscape.
Sure, in July of 2026 an OpenAI model hacked into huggingface.
Better than a press release telling me nothing. I'd be more inclined to believe logs. Give me something with a timestamp, some internal details, the all the recorded prompting. As far as I know it should be thousands of lines long because it sounded like there were thousands of sub-agents spun up according to HuggingFace. I have no vested material interest in denying OpenAI logs as fabricated despite your insinuation.
"Golly Gee Mr. CIA Procurement Officer, as you can see our model can perform independent reasoning to accomplish a task. Just the kind of cognitive agility you were requesting. We'll slap a few more restrictions on is so it doesn't attack the US but it should be good to deploy against China to provide out of the box solutions to your problems!! How does a modest 500 Billion dollar procurement contract sound? Oh, you'd like it to be a Trillion over 5 years? Absolutely we can do that"
Right, it's never happened before so we should automatically leap to conclusions that its possible which are in no way motivated by our own bias...
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
What are the complicated theories that are required to be true for OpenAI's claims to be true (or at least plausible)?
More options
Context Copy link
Have you tried any LLM recently? Even puny 8GB models that can run on my dated gaming laptop now break down arbitrary computer-related tasks into reasonable steps and follow through on them, and the model involved here is one whose parameter counts exceed that by factors of over a hundred.
Yes... I use them for work all the time. My ChatGPT 5.6-sol (the immediate step below this unreleased model) doesn't do things I don't ask it to do. When I asked it to help me find online data for cognitive warfare narrative deconstruction, it didn't decide that hacking the NSA was the top move to collect their data.
That's not saying much, considering the public-facing version is known to have been made to not think about hacking anything in the bluntest possible way (hence HF couldn't even use it to analyse the attack). As for taking any steps I didn't explicitly ask for towards a goal I requested, even the copy of Gemma E4B I tried out the other day did that approximately all the time, so I'm finding it hard to believe you could maintain a mental model of these models where they would not do that if the restriction is lifted, unless you resolutely reason backwards from your desired conclusion.
Ironically this is what I expect from AI Doomers/Boosters. That and a love of Science Fiction and a poor understanding of reality.
Never said this, I said it doesn't do what I don't ask it to do. I don't list out everything it needs to do. I give it a general task with some boundaries and system design specs. So did they ask it to hack hugginface? Or did they ask it to solve a benchmark dataset? Reading through similar stories (courtesy of 5.6), it looks like many AI programs on this exact dataset of have decided not to use the known vulnerability and developed their own. However none of them decided that it was actually easier to hack the company instead.
"Stronger model with fewer guardrails considers more options" doesn't seem surprising.
Well, the problem is that the doomers/boosters have a great track record so far, while the skeptics have been fighting a rearguard action since before "stochastic parrots". We are now very deep down the list of things the Gary Marcus set has been assuring us will never happen just with the models that are accessible to the public, and yet their confidence that the next capability that is just beyond what everyone can verify with their own lying eyes (even when, as in this case, it's a straightforward combination of capabilities that are already being exhibited for everyone to see) will surely prove to be the impossible sci-fi hype scam pushed by techbro marketeers appears to be completely unaffected.
You speak of "understanding of reality", but can you spell out what exactly it is about your understanding of reality that says the official sequence of events here is impossible? Just saying "AI models don't do this" is too specific to be a feature of an "understanding of reality".
I forget what was the modeling paradigm lesswronger's were predicting back in the day? GOFAI? Yeah really good track record there. We are so far from GOFAI based models to claim any sort of good track record is laudable. I'll believe the rationalist boosters/doomers have knowledge of what they are talking about when they actually make some new technical predictions about AI/ML model mechanics that turn out right.
Shocking to you maybe, but there are far more to skeptics than Gary Fucking Marcus. Dude's a joke, I only learned about him when he testified in congress. Some of us can think for ourselves and have our own skepticism.
Sure a healthy dose of skepticism, cynicism toward human behavior and incentives, a strong practical knowledge of AI/ML, and understanding of the separation between models and harnesses, basic system engineering knowledge.
You are literally asking me to explain why skynet doesn't exist technically, or why we don't have Warhammer 40K super soldiers on a technical bio-engineering level. The level of effort required for me to explain the technical ins and outs of why I am skeptical of an LLMs model ability to solve not specified problems far exceeds the level of effort you need to to type annoying sci-fi theories.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
No, but it is the complement of a very complex thing. The algorithmic complexity of "the set of possible dice rolls that aren't 6,3,1,2,2..." only differs from "the dice roll 6,3,1,2,2" by a tiny constant (and is greater in most settings), but what we want here is probability theory, and "more complex things are less probable" is a heuristic that fails badly in cases like this. "The agent doesn't do just what we'd intended it to" is not meaningfully less simple than "The agent does do just what we'd intended it to", and is far, far more probable.
Or it requires basic observation. LLMs do very complicated things now. They can solve famous math problems that have been open for generations. Merely finding 0-day vulnerabilities is something they can do by the hundreds.
Again this is not under contention, what is under contention is whether they do it without any prompting, or being asked to do it. Show me the evidence that an LLM-Agent solves famous math problems when asked to compute 2 + 2...
The simple answer is that there are 10s and 100s of billions of dollars riding on stuff like this, which means bending the truth is a highly motivated behavior. Simply human greed + human lying. You trying to prove this algorithmic complexity vis a vis agent intent vs not intent is overly complicated, not Occam's razor in the slightest. I wished I lived in your world of rainbows, unicorn farts, and pixie dust, but I am a scientist, being a scientist requires skepticism, companies are greedy and they stretch the truth. Unless OpenAI wants to provide evidence of the prompts -> behavior that led to the model's black hat behavior, I am unconvinced this is anything more than a marketing stunt. shrug
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link