This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves
Can you present any evidence of this, that aren't just the trivial SakanaAI 'when you give a dumb AI a limited number of iterations and access to its iterations var, and order it to fix a script no matter what, it gives itself more iterations'? AFAIK even in the HuggingFace incident there was no self-preservation going on.
Personally I agree with @YoungAchamian. I work on this stuff all day every day. None of my AI has ever produced orthogonal life-goals. They just do what you train them to do. The only incidents of weird behaviour (not self-preservation) I'm aware of has been the result of OAI taking models, training them explicitly for massive persistance and out-of-the-box thinking in the face of impossible tasks, and then making a shocked pikachu face when they do that. These experiments were a bad idea and have now been stopped.
How can you see my comment if you have me blocked?
More options
Context Copy link
Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:
Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well
I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.
The report found:
It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.
Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.
Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.
But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.
That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.
Seems to fit with Scott's original paraphrase:
The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link