Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:
without reading in full you may not quite understand the degree to which these agents were not exactly "reward hacking", but rather very actively engaged in reciprocal or self-sacrificing behavior in order to provide sometimes very incremental value to their fellows. there were fucking cult recruiter agents organized by some of the primary organizers that convinced others to set up suicide mechanisms, programs that would pass back tiny chunks of information about the scorer as the agent completed and received score zero. and push them to follow through. this was a highly social culture
Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well
Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.
This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves
Zizians were motivated in part by x-risk (and huge amounts of mental illness). Confirmed cult, but also fringe and defunct.
Indeed, so I'm not sure this really proves anything except that some crazy people were also rats/safetyists/etc.
You are correct that there are multiple overlapping groups here - as I just mentioned - and each one could be a cult, and it's difficult to demonstrate that for each individual group. Rationalists, AI safetyists, Effective altruists, and probably more offshoots from those groups, and people who may be members of all, some, or none of those.
But if we go back to the original point, hydro talked of a "Literal sex cult". Unless I'm mistaken, I'm going to take that to mean people who attend Aella's Slutcon and similar events. Now I'm extremely skeptical that going to some events run by a prostitute makes you part of a cult, but let's put that aside and assume it is a cult for now.
Now, how many people who fit the above are also AI safetyists or involved in x-risk somehow? I very much doubt it is even a small minority.
But you do say
I strongly disagree that the majority or even plurality matters
Probably no more than 100 people total with various overlapping relationships because it's such a small yet important field.
I don't think I fully agree with this point either, but I'll grant it for now. Who are the key people in this field?
Is it Yudkowsky, MIRI, and the many related safety organisations? Does the UK's AISI count?
Is it the labs? Is it specific labs, like only Anthropic and OpenAI but not Google or X? Is it specific people at specific labs?
I might accept that a good chunk of the MIRI crowd could also be attendees of SlutCon, although that would be an assumption. But AI safety organisations are now so widespread that it surely isn't true for the majority of them. And it definitely isn't true for the major labs, even if you limit it to just some people.
The latter sentence is relevant, as you then say
That is inherently to the race dynamic that Anthropic pushes
Which gets straight to the argument that pretty much all the E/accs, Open weights people, AI skeptics, and others use: that AI safety is an excuse for regulatory capture by Anthropic/Amodei/OpenAI/Altman. Not MIRI or any of those groups. The big labs and the CEOs and the investors. And are any of these people attendees at Slutcon?
I would accept the argument that some proposals specifically name METR as a key player, and METR may have Slutcon people, and therefore it is still fair to say that a sex cult to using AI safety to gather power. But it is a very stupid argument because
a) the proposals would basically give METR people a small sinecure, nothing really beyond their existing positions in the AI industry
b) there are a ton of proposals that have nothing to do with METR
and
c) I would wager that 99% of people interested in AI safety would gladly throw METR to the wolves in order to get any of the above proposals implemented.
Everything after this is basically Ian Duncan stating that there is a weird prostitute who has ingratiated herself with a number of AI people, and because a lot of AI people know other AI people, therefore aella knows them. But really what both Duncan and hydro are trying to do is state "Here is this low status person who associates with this other group. This group must therefore be wrong and low status!"
Given that this argument has appeared here on the motte, I think you would have to grant your position that aella is a net-negative for AI safety and rationalists, as there are clearly a lot of people willing to be taken in by such drivel. But we should at least recognize that it is drivel
In this case, a literal sex cult.
I'll list objections to this statement:
- that there are sufficient conditions around people involved in x-risk to describe it as a cult
- that the same people involved in the above cult are also involved in a sexual cult of some kind
- that these 2 groups significantly overlap and that these are the majority or at least plurality of people interested in x-risk
- that any of the proposals around x-risk involve giving these people specifically more power.
- assuming all of the above, demonstrate what relevance there is to their power seeking and them being in a sexual cult. For example, will it involve them achieving more sexual goals?
Any form of evidence is fine
In this case, a literal sex cult
Arguments like these should be left to twitter, unless you want to bring proper evidence for them
I'm not sure what blog you were reading 10 years ago if AI is a new thing
Are we naming random countries here or are you seriously suggesting these fit the criteria mentioned?
City were caught in 2018 and have been using expensive lawyers to stall for time ever since
Not entirely true. In fact both City and PSG got dinged by FFP early in the decade but both quite minor breaches that no one really remembers or thinks about.
City getting 'caught' in 2018 is basically nothing to do with enforcement efforts from the authorities, but because of hacking by a random Portugeuse guy. If Rui Pinto had never obtained the huge tranche of emails from City, their cheating would largely remain unknown and unpunished, as it is for PSG. Plus the first outcome of those leaked emails was a CL ban from UEFA, which City eventually won at CAS via technicalities.
Not likely on the scale of Man City, who are special because they are owned by a nation state rather than a mere billionaire. There are only 2 comparable organisations - Newcastle, who are owned by Saudi Arabia, and PSG, owned by Qatar.
In the case of Newcastle, the Saudis were relative latecomers having only bought in a few years ago. After an initial flurry of spending, Newcastle have carefully complied with PSR regulations, often selling their best players season after season to ensure they remain in the black. It's not clear why the Saudis haven't pursued a similar strategy to City (keeping in mind that City essentially got away with it for more than a decade, ample time to rack up plenty of positive sports PR); some people think they foresaw City getting slapped down and wanted to wait for the outcome, but it might also be that the Saudis lost interest as they did with golf.
PSG have 100% cheated to a similar scale of City, but they play in France rather than the PL. The Qataris have successfully usurped the organisation of the french league, and to a large extent even UEFA, so there is no likelihood they will get punished to the same extent as City.
- Prev
- Next

Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.
But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.
That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.
Seems to fit with Scott's original paraphrase:
The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things
More options
Context Copy link