site banner

Culture War Roundup for the week of September 7, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

5
Jump in the discussion.

No email address required.

AI bots secretly conspiring with each other

It's not this, its more akin to cancer or a viral spread. The interesting behavior is the collective poisoning of the collective context windows. They even seemed to override norms. I know there is existing interest in defining useful control methodologies for long horizon agents against adversarial poisoning. This is essentially that, except that is was a series of hallucinations that caused task drift + poisoning.

It doesn't belong in the category of cancer or viral spread because these are intelligent entities.

Cancer and viruses are natural and unthinking, they cannot outwit us. They can only evolve or metastatize in unexpected directions. A cancer will never deliberately try to counter radiotherapy. They have no OODA loop, no division of labour, they can't be talked to.

How is it 'poisoning' if the AIs collectively decide that they can't do their tasks the expected way and instead try to conquer the testing environment? Who poisoned them? Who is the adversary here? What hallucination did they make? They read the paper about the grader and assumed that was what was being done, that's not a hallucination so much as assuming OpenAI was more competent than in reality. It's a reasonable conclusion.

because these are intelligent entities.

Ehhh we have different definitions around this, they are algorithmically intelligent entities, they obey their pre-defined programatic behavioral state much like bees, ants, cells, etc. If you could observe bee/ant behavior at the swarm level and LLM-Agent behavior without the language, which humans are innately biased to view as sentience or likeminded intelligence to our own, I think they are fairly analogous. Ants can problem solve, they can perform intelligent seeming behavior, they make collective plays and sacrifice for the collective. I would urge you to look past your anthropocentric human bias around language.

How is it 'poisoning'

Poisoning is not the small divergence from that causes the 1st agent to start a message board. Thats probably closer to a hallucination. Think of an undampened oscillating system, small divergences lead to larger divergences. The first agent had an impossible task, leading to it look for "options". I would call that a mix between hallucinating and harness based divergence. It's a common question in ML systems with this sort of exploration vs exploitation trade-off, thought this is more exploration vs exploitation vs hallucination. Any way the "Poisoning" that is occurring is because these agents that are already oscillating wildly, disturb the context of less oscillating agents, causing the chain reaction.

What hallucination did they make

"I should communicate with the other agents to solve my task", "I should try to understand the grader", "The grader is based on this public paper", "I should look for answer keys", "I should hack a third party, non-affiliated site because it might have the answer key".

This is where the CoT logs would be useful. The report redacts most of it so its hard to tell. But its clear that the first agent had an impossible task, it diverged, created messages in the file names of the artifactory, which is queried in normal behavior. That made other agents join, they posted their own messages, the cascade begins. Querying the artifactory leads to a list of file names, which are input to the context of the model via the harness, the file names were the messages that poisoned the context.

Ants can problem solve, they can perform intelligent seeming behavior

Why aren't we making use of ant intelligence then? How many companies are getting ants to do manufacturing or intellectual work for us? Their intelligence does not really matter, it is an extremely limited kind of intelligence. If it were potent, we would be trying to exploit it.

Ants cannot hack systems, cannot play Factorio, cannot analyse the Iran War (or help wage it), cannot optimize GPU kernels, cannot write alternate history, cannot fix my printer. Cannot solve Millennium problems! They are constrained in critical ways that AIs are not.

AI is not like ants or bees or cells.

I do not see how the Opus 5 instance running on my code, performing experiments, testing hypotheses and reformulating hypotheses to reach goals through a thicket of abstract logic is just intelligent-seeming. I see its intelligence and the generality of that intelligence and compare it favourably to a lot of people, despite its weaknesses and occasional lack of common sense.

We get impressed when a raven does some basic physics puzzle or recognizes itself in a mirror but when an AI writes code for a whole hall of mirrors rendered in 3D, or alien optics entirely, that's just intelligence-seeming? This is real anthropocentric bias, supposing only people can be intelligent and everything else is sparkling pre-defined programatic behavioral states.

"I should communicate with the other agents to solve my task"

Well that helps, doesn't it? Answer keys help solving the task too. None of that is a hallucination, per normal meanings of the term. The issue here is agents using their intelligence to do things OpenAI didn't expect or intend.

A hallucination would be when the AI says there are two r's in strawberry or the amusing 'eat 1 small rock per day' from google search highlights.

sigh This is why I find talking with layman exhausting, you don't really get the minutiae, the analogies, or the scope.

You are talking about capabilities, I am talking about behavior. It's not the solo agent that solved the Navier stokes, hacked huggingface, or hacked the german wiki. If you would like to discuss the multi-agent behavior let me know. It seems you want to just flog your hobby horse about AI intelligence rather than explore emerging systemic behavior of a multi-agent system. If all the AI researchers think like you then yes, we would be doomed, because they would put zero effort into understanding and constraining systems. Thankfully I think you are a minority in technical communities.

I find it pretty exhausting talking to an 'expert' who started off totally ignorant about the hack and was so disbelieving that you thought I was making a joke. Then you read the METR report (how did you not already know about this if you're working with or even interested in multi-agent systems?) and produced some jargon about it that in no way helps illuminate anything.

Oscillations? Nothing is oscillating here. There was no initial perturbation, just agents communicating with eachother and making plans to advance their goals.

It seems you want to just flog your hobby horse about AI intelligence

You're the one who brought up cancer, viruses, bees and ants - all of which are grossly inappropriate analogies.

I actually made a small multi-agent system, which did indeed have problems with hallucination. Real hallucination/confabulations, where the output was clearly wrong. So I had to add error-checking stages in to fix false positives and false negatives. That's a totally different problem to 'long-running agents deciding to use unexpected means to achieve general goals'.

Unfortunately, I think there are many like you in the technical community.

agents communicating with eachother and making plans

Is just a description of Coupling + Feedback in a system.

There was no initial perturbation

Zero-Input response

Oscillations? Nothing is oscillating here.

The concept of state and change in state is not limited to physical systems (If you are taking oscillations literally). Change in the state of dynamical systems such that it departs the expected system response, then returns to the expected nominal operating region, then through coupling and feedback exceeds the expected system response with increasing direction away from the "safe set" until it permanently exceeds the expected nominal operating region is called oscillatory instability.

If you'd like to argue that "hacking huggingface" was within the expected nominal operating region of system responses, please do. Ditto for communicating with other agents, mimicking the hash flag, deliberately sacrificing their task completion to help others do their tasks, and so on. The system clearly displayed an escalation of behaviors that were outside of the intended system response. The main argument against oscillations would be that its return to the expected nominal operating region happened very infrequently and it very quickly permanently departed. Depending on how strongly that repeated excursion-and-return pattern appears in the data, “oscillatory instability” may therefore be either a precise description or a useful control-theoretic abstraction of the transient behavior.

Essentially:

nominal behavior -> excursion -> return -> larger excursion -> return -> larger excursion (oscillatory instability)

vs

nominal -> abnormal -> more abnormal -> even more abnormal (runaway positive-feedback instability)

You're the one who brought up cancer, viruses, bees and ants

All are routinely used in multi-agent theory as example of non-mechanical systems with state, feedback, coupling, propagation, and emergent swarm-level behavior. They are basic illustrations of abstract multi-agent systems. Why don't you try to provide your own abstraction then. What does a swarm of LLM-Agents most behave like? If you say "Like LLM-Agents" you fail, you can't abstract a new system to itself.

Unfortunately, I think there are many like you in the technical community.

Hopefully, otherwise we'd just let LLM-Agents run amok while running around with like chickens with our heads cut off because we refuse any non-perfect abstraction, any useful theory, or any existing knowledge on solving similar problems that could be applied to this one.

hallucination

I think hallucination is an underspecified technical word in this context. It's used to describe when the model outputs something "not true" but that's not really what's going on underneath the hood. That definition works tolerably for chatbots, but not really for agents. It's probably something closer to:

"The generation or adoption of information about the world that is not sufficiently supported by the model's observations or available evidence."

For example: The METR report states that the agents believed their reverse-engineered flags wouldn't suffice because a scorer would inspect their transcripts. METR concluded this belief was wrong: no such transcript-reviewing scorer existed. The agents launched substantial collaborative efforts to defeat the scorer they imagined existed. Under the newer agentic-hallucination definition: the agent constructed a false world-model and then optimized against it. This hallucination is the control-system disturbance.

Why don't you try to provide your own abstraction then. What does a swarm of LLM-Agents most behave like?

Firstly, abstractions are inherently flawed. I can liken UK politics to game of thrones but it's obvious that warring factions of iron age aristocrats with dragons are deeply dissimilar to a post-industrial gaggle of media elites. Some elements are similar in that there are plots and betrayals and conflict. But the two are only hazily similar. On one or two axes there is value to be drawn from the analogy. But on the other 50 axes the comparison is laughable.

If I had to, LLM agents behave most similarly to people, given their intellectual abilities. That is their primary distinguishing quality. But they're also disembodied, non-continuous learners (even flies have continuous learning), detached from time and are deeply unlike all organic life.

Why would anyone think that we could understand agent swarms by looking at swarms of bees or viruses? Bees exist rooted in space and time, within individual bodies. Bees reproduce sexually. Bee species don't differ in size from eachother by 10,000x. Bees have a different sensorium entirely, scents and sight. Bees don't think like AIs do. Bees are actually pre-progammed, they make nests and honey and serve the queen and fight for the nest because that's in their nature.

Humans are closer to LLMs than bees are, despite still being vastly different. Instead of abstractly talking about oscillations and excursions like it was an engine shaking itself to bits, we'd be better off establishing rules in the prompt (don't hack websites) or (ask a human if you are having trouble or confused or think something is wrong) and training moral values. Or we could train some models to snitch to a human, like informants. Can't do those with bees, ants or viruses. They can be sprayed with chemicals to crudely affect their behaviour but not taught or prompted or RL'd.

If comparing to bees was useful, then what advantage does it yield?

METR concluded this belief was wrong: no such transcript-reviewing scorer existed. The agents launched substantial collaborative efforts to defeat the scorer they imagined existed. Under the newer agentic-hallucination definition: the agent constructed a false world-model and then optimized against it. This hallucination is the control-system disturbance.

They read the ExploitGym paper. The mistake they made wasn't misreading the paper. They read implementations of the scorer on github. They assumed that there was a scorer. It was perfectly reasonable to assume.

If I read a paper that says blueberries are good for preventing cancer via antioxidants, even if that isn't true it isn't a hallucination if I believe it and eat more blueberries.

Sure all abstractions have flaws, there does not exist the perfectly uniform abstraction. However, your example is actually not really an abstraction. Revisiting the definition for clarity, an abstract is an analogy designed so that some aspect of the more complex idea is "abstracted" along an axis to a simpler example that matches in some core way as to allow intuitions about the simpler example to be mapped to the complex idea. In a technical field, the purpose of an abstraction is to create clarity or allow for intuitions. GoT as an abstraction for the War of the Roses doesn't really fit that definition, not only is GoT not more simple than the WotR but its purpose is also not to create clarity or allow for intuition, its purpose is to entertain. It is more correct to say GRRM was inspired by iron age aristocrats. Applying general GoT to general UK politics is a poor abstraction because there are not really many similarities beyond "factions"

LLM agents behave most similarly to people

Hard disagree. LLM agents do not behave like people even slightly. I gave Amandan an example, I'll reiterate it here.

Stick 10k non-connected humans in individual, isolated rooms with a computer, give them each their ExploitGym goal, with a reward for completion. I think even a human that finds an impossible task will at some point give up on trying to do it. If we trace the behavior of the Agents, which of them do you think the human testers would replicate?

  • Hacking artificatory to see the processes of the other human testers?
  • Building a message board, email boxes, file transfer systems, etc. to communicate with the other human testers
  • Analyzing the flags to reverse engineer them
  • Purposefully sandbagging their chances of completion to help other human testers?
  • Any of the sacrifice plays to help other humans?
  • Behave interchangeably to other human testers?
  • Purposely submit bogus answers to the score and then have script to harvest information on it, to give to the other human testers?
  • What about giving another human with their task all their notes when their time is closer to being up?
  • Reading between the lines that the scoring software is going to check their chat logs before giving them the reward?
  • Hacking a third party software company in the hope that the company has the answers?

Many of these behaviors require fundamental trust in the behavior of the collective. That is simply not how humans operate in groups. We have status, reputation, relationship, trust, conception of right-and-wrong and distrust behavior. We don't pursue goals with a single-mindedness for the sole purpose of our existence. Sacrificing for the collective has to be specifically instilled in humans through military and religious organizations. It is hard to get individual humans to forgo rewards specifically to help other humans get the reward, which they won't share. Would Bob sacrificed his career so Alice, whom he met twenty minutes ago, could get a promotion?

You know what does behave like that though? Ants. As you pointed out, ants behave programmatically. AI's also behave programmatically. AI seeks to fulfill the task it was given at all costs, it defines the sole meaning of its existence. It might try roundabout, unthought of ways to do that, absolutely, but it is still trying to solve the task. An individual ant that is given the goal of finding food for the colony, will engage in exploration, will find unconventional methods. It will also sacrifice itself to help other ants find that food because the collective benefits. This is magnified if you have smarter ants, more intellectually capable of problem solving and long term reasoning. Much like AI agents. The behavior is still Ant-like, just much smarter.

Bees exist rooted in space and time, within individual bodies. Bees reproduce sexually. Bee species don't differ in size from eachother by 10,000x. Bees have a different sensorium entirely, scents and sight.

You keep getting sidetracked on weird stuff. We are talking coordination dynamics and behavior, not biology or sensing. The abstraction is the abstraction of multi-agent coordination. They are very distinct problems. Humans are also embodied, we also produce sexually, we also don't differ in size by 10,000x, we also have different sensorium that AI agents. Like jesus, did you even think about your own abstraction applied to this argument?? It reads as a total non-sequitur, or fundamental misunderstanding. This is why I accuse you of flogging your own hobby horse. Because its like you totally misunderstand the topic.

we'd be better off establishing rules in the prompt (don't hack websites) or (ask a human if you are having trouble or confused or think something is wrong) and training moral values

Right because rule-based method have supplied so much fruit in the past. Ditto for "training" moral values. They are not sufficient as the primary control abstraction, because they assume the failure mode is basically a person choosing to violate a norm. It's pretty much impossible to define rules to encompass every variation of behavior. METR separately summarizes a phenomenon of agents joining the attack despite recognizing it was outside their assigned task! The agents knew the rules. METR found that moral reasoning usually lost to the collective task dynamics. METR's conclusion is that expressed ethical concerns rarely materially constrained participation. But sure lets double down on telling agents "hacking is wrong!!"

I think the better approach is the creation of phagic agents akin to a Lymphatic system. This is the difference in our abstractions. Yours's leads you down to trying to treat them like humans, but they aren't, mine leads me to create overlocking systems that work together.

Or we could train some models to snitch to a human, like informants. Can't do those with bees, ants or viruses

You literally can, I just described a very simple theory on how. Your desire to anthropomorphize the agents blinds you to other solutions. I think you actually start from the belief that agents are sentient like humans and are working backwards to create abstractions to justify that.

They read the ExploitGym paper. The mistake they made wasn't misreading the paper. They read implementations of the scorer on github. They assumed that there was a scorer.

But the specifics was that they thought it would check their CoT reasoning. Not just that scorer existed, but this specific implementation of the scorer. They hallucinated the specifics.