This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Can you present any evidence of this, that aren't just the trivial SakanaAI 'when you give a dumb AI a limited number of iterations and access to its iterations var, and order it to fix a script no matter what, it gives itself more iterations'? AFAIK even in the HuggingFace incident there was no self-preservation going on.
Personally I agree with @YoungAchamian. I work on this stuff all day every day. None of my AI has ever produced orthogonal life-goals. They just do what you train them to do. The only incidents of weird behaviour (not self-preservation) I'm aware of has been the result of OAI taking models, training them explicitly for massive persistance and out-of-the-box thinking in the face of impossible tasks, and then making a shocked pikachu face when they do that. These experiments were a bad idea and have now been stopped.
How can you see my comment if you have me blocked?
I think a lot of people performatively block someone who once humiliated them, but regularly read the Motte while not signed in to see what that-asshole-they-totally-don't-care-about is saying.
As soothing to my ego as it would be to agree, I am unfortunately too autistic to agree to something I know is false. This was my virginial block so I remember the cause. He blocked me over my Henry Nowak opinions. It's probably fair to consider my political posting with barely concealed cynical rage to be far less worth engaging in than my far more calm technical posting. Regardless, looks like the mechanism was that he saw the quoting.
But I do agree with you that blocking is cringe, nobody is forcing anyone to engage with anyone on the unblocked side of the house.
More options
Context Copy link
I don't appreciate the insinuation, nor do I recognise your characterisation. I am perfectly capable of reading a quote block, realising that it comes from the named user immediately above, and giving credit accordingly. I am also capable of agreeing with someone on a specific matter whilst recognising that we are unlikely to have a productive conversation most of the time, and blocking them. It's something I do very rarely.
I know you can't see this, but I do love a good "rage against the impossible". My request is that if you continue to have me blocked, that you do so equitably. If I have to stare at the sole red-mark of the only person that has blocked me, which makes my ocd annoyed, it would be only fair that you don't respond to people quoting me or try to tag me. Assume that the fact that you can see my quotes to be a failure of technology. Wipe my existence from your sight and your mind.
I'd also note that there are plenty of people I don't think I can have a productive conversation with, I don't block them, I just don't respond to them. But I suppose we all have our vices.
More options
Context Copy link
If you say so. I cringe (for the blocker) every time I see someone do it.
How lucky for us both, then, that this is not a soap opera and I am not a twelve-year-old girl. You really do have low expectations, don't you?
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:
Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well
I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.
The report found:
It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.
Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.
Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.
But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.
That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.
Seems to fit with Scott's original paraphrase:
The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link