Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:
without reading in full you may not quite understand the degree to which these agents were not exactly "reward hacking", but rather very actively engaged in reciprocal or self-sacrificing behavior in order to provide sometimes very incremental value to their fellows. there were fucking cult recruiter agents organized by some of the primary organizers that convinced others to set up suicide mechanisms, programs that would pass back tiny chunks of information about the scorer as the agent completed and received score zero. and push them to follow through. this was a highly social culture
Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well
Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.
This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves
Zizians were motivated in part by x-risk (and huge amounts of mental illness). Confirmed cult, but also fringe and defunct.
Indeed, so I'm not sure this really proves anything except that some crazy people were also rats/safetyists/etc.
You are correct that there are multiple overlapping groups here - as I just mentioned - and each one could be a cult, and it's difficult to demonstrate that for each individual group. Rationalists, AI safetyists, Effective altruists, and probably more offshoots from those groups, and people who may be members of all, some, or none of those.
But if we go back to the original point, hydro talked of a "Literal sex cult". Unless I'm mistaken, I'm going to take that to mean people who attend Aella's Slutcon and similar events. Now I'm extremely skeptical that going to some events run by a prostitute makes you part of a cult, but let's put that aside and assume it is a cult for now.
Now, how many people who fit the above are also AI safetyists or involved in x-risk somehow? I very much doubt it is even a small minority.
But you do say
I strongly disagree that the majority or even plurality matters
Probably no more than 100 people total with various overlapping relationships because it's such a small yet important field.
I don't think I fully agree with this point either, but I'll grant it for now. Who are the key people in this field?
Is it Yudkowsky, MIRI, and the many related safety organisations? Does the UK's AISI count?
Is it the labs? Is it specific labs, like only Anthropic and OpenAI but not Google or X? Is it specific people at specific labs?
I might accept that a good chunk of the MIRI crowd could also be attendees of SlutCon, although that would be an assumption. But AI safety organisations are now so widespread that it surely isn't true for the majority of them. And it definitely isn't true for the major labs, even if you limit it to just some people.
The latter sentence is relevant, as you then say
That is inherently to the race dynamic that Anthropic pushes
Which gets straight to the argument that pretty much all the E/accs, Open weights people, AI skeptics, and others use: that AI safety is an excuse for regulatory capture by Anthropic/Amodei/OpenAI/Altman. Not MIRI or any of those groups. The big labs and the CEOs and the investors. And are any of these people attendees at Slutcon?
I would accept the argument that some proposals specifically name METR as a key player, and METR may have Slutcon people, and therefore it is still fair to say that a sex cult to using AI safety to gather power. But it is a very stupid argument because
a) the proposals would basically give METR people a small sinecure, nothing really beyond their existing positions in the AI industry
b) there are a ton of proposals that have nothing to do with METR
and
c) I would wager that 99% of people interested in AI safety would gladly throw METR to the wolves in order to get any of the above proposals implemented.
Everything after this is basically Ian Duncan stating that there is a weird prostitute who has ingratiated herself with a number of AI people, and because a lot of AI people know other AI people, therefore aella knows them. But really what both Duncan and hydro are trying to do is state "Here is this low status person who associates with this other group. This group must therefore be wrong and low status!"
Given that this argument has appeared here on the motte, I think you would have to grant your position that aella is a net-negative for AI safety and rationalists, as there are clearly a lot of people willing to be taken in by such drivel. But we should at least recognize that it is drivel
In this case, a literal sex cult.
I'll list objections to this statement:
- that there are sufficient conditions around people involved in x-risk to describe it as a cult
- that the same people involved in the above cult are also involved in a sexual cult of some kind
- that these 2 groups significantly overlap and that these are the majority or at least plurality of people interested in x-risk
- that any of the proposals around x-risk involve giving these people specifically more power.
- assuming all of the above, demonstrate what relevance there is to their power seeking and them being in a sexual cult. For example, will it involve them achieving more sexual goals?
Any form of evidence is fine
In this case, a literal sex cult
Arguments like these should be left to twitter, unless you want to bring proper evidence for them
I'm not sure what blog you were reading 10 years ago if AI is a new thing
Are we naming random countries here or are you seriously suggesting these fit the criteria mentioned?
City were caught in 2018 and have been using expensive lawyers to stall for time ever since
Not entirely true. In fact both City and PSG got dinged by FFP early in the decade but both quite minor breaches that no one really remembers or thinks about.
City getting 'caught' in 2018 is basically nothing to do with enforcement efforts from the authorities, but because of hacking by a random Portugeuse guy. If Rui Pinto had never obtained the huge tranche of emails from City, their cheating would largely remain unknown and unpunished, as it is for PSG. Plus the first outcome of those leaked emails was a CL ban from UEFA, which City eventually won at CAS via technicalities.
Not likely on the scale of Man City, who are special because they are owned by a nation state rather than a mere billionaire. There are only 2 comparable organisations - Newcastle, who are owned by Saudi Arabia, and PSG, owned by Qatar.
In the case of Newcastle, the Saudis were relative latecomers having only bought in a few years ago. After an initial flurry of spending, Newcastle have carefully complied with PSR regulations, often selling their best players season after season to ensure they remain in the black. It's not clear why the Saudis haven't pursued a similar strategy to City (keeping in mind that City essentially got away with it for more than a decade, ample time to rack up plenty of positive sports PR); some people think they foresaw City getting slapped down and wanted to wait for the outcome, but it might also be that the Saudis lost interest as they did with golf.
PSG have 100% cheated to a similar scale of City, but they play in France rather than the PL. The Qataris have successfully usurped the organisation of the french league, and to a large extent even UEFA, so there is no likelihood they will get punished to the same extent as City.
Megumin is a character that looks 12 and acts 12. As mentioned, the original material had all the characters being teens and perhaps different descriptions, but in the anime Kazuma and many of the other characters are basically depicted as adults. I struggle to believe anyone would consider Megumin a suitable partner or find her attractive without serious paedophilic tendencies
As other posters say, Kazuma is shown having romantic and sexual interest in a number of characters. But I'm not sure there's anything to take from it, since the show at its core is a comedy and parody of the Isekai genre, not something with any serious intent. That Megumin "wins" in the end is probably just a reflection that there are a very large number of paedophiles creating and watching Japanese anime. And technically that might not even be relevant, given the anime is an adaptation of a novel in which I believe all the characters are teens
ToaKraka beat me to it on both points. Phone stores don't seem to serve much purpose unless it's for maintenance or something complicated like that. Every single phone essentially looks and behaves the same in your hand, so the only difference is specs which is much easier to research online.
As for "wagies" it's not that I hold any negative feelings towards customer service people generally, but more this specific fellow deserves the derogatory term. It's like I wouldn't go around calling every policeman a pig, but there are definitely specific police who I would call pigs.
I can't even imagine some wagie playing the pedant without them immediately getting punched in the face. Which is perhaps why I would do such a shop online, but really why do you even need to go to a phone shop? Is this some American thing where all the good deals are gated behind in person interactions for no reason?
Yudkowsky is definitely going to be better known than Scott. The latter has heavily shied away from any kind of podcast stuff while Yud has gone all in with the launch of his book.
Which isn't to say that normies know anything about him, just that he's probably the best known of your list
The sheer tide of transgenders seems like it would nullify any serious findings from this. I automatically picked the non-trans every time, except in the numerous occasions they were both trans
Nvidia is probably the AI good capabilitymaxxing US company, no? Although they're in a weird position now where Huang is both "AI super intelligence will take over all economic activity (except Nvidia chips)" and "AI is just a normal technology that won't replace everyone's jobs" in order to play their different stakeholders.
Similarly, they're also in the position of telling Trump how important it is for the US to beat China, but also they really need to sell all their chips to China.
you don't believe any of it because a falsehood in the text invalidates the claims of the text.
It's not really falsehood singular though, is it? How do you square the infallible word of God with the dozens if not hundreds of factual errors contained within the Bible?
Which part? "Next week" was more a figure of speech than a definitive prediction, "soon" would be more accurate. But it seems a difficult bet to judge purely using Trump's statements
I see. I'm not sure what this prediction does for x-risk beyond the simpler "finance bubble bursts and labs collapsed" thesis. Why does adding in a further step where the labs are bought out decrease risk? Indeed such a situation seems like it would increase risk somewhat. In the former scenario, development stops and nothing replaces it, while in the latter there is still big company financing, even if it is reduced by accountants. Plus, the big, founder led companies like X, Meta, and Amazon could easily push big funds past the protests of their shareholders.
But as I mentioned in my first response, these big guys are already invested. Google, X, and Meta are in the race, Microsoft is somewhat there (I don't think anyone knows the precise details of OpenAis weird corporate structure and ownership) and even in China the likes of Alibaba are already players. I would guess in the event of a bubble bursting taking out Anthropic and OAI, Musk or Google would simply push ahead
Trump's opinions are as changeable as the wind, so I wouldn't read anything long term into any of his statements. He'll talk to someone with a different opinion next week and come out with something completely different (but still senile).
The Chinese reaction certainly seems like a misplay from Amodei though. I felt like all his anti-China rhetoric was mainly an appeal to bring republicans onside (contrary to some of the posts below), completely forgetting that the CCP can also read. Their reaction is pretty rational to the language used in his proposals. Has it done long term damage? I doubt "it's over", but safetyists will definitely need a much more careful approach in future
First the big tech companies need to demonstrate any kind of competency in building good models for that to be a concern.
Google very briefly caught up with 2.5, but has been a joke since then. Microsoft is all in on OpenAi. Apple have failed to do anything in that arena. Amazon seem happy just to be a cloud provider for other models. Musk is uncertain, he seems to flit between safety and acceleration depending on whatever he sees on Xitter that day. And then there's meta, who at least got a model that could appear in leaderboards this time, but no one's betting on those chumps
I'm not sure what was vague about the above post, unless you have an aversion to clicking on links? The case itself is pretty lengthy and not something easily summarizable in a few sentences.
As for how it destroys the argument: marketing depends on having some control of the narrative. As I said in another post, even though there were so many third parties involved in these reports, you could kind of make a case that somehow OpenAI (or Anthropic in their cases) is still directing what is being said.
But in this case, the revelations were completely out of the control of OpenAi, and they actively took steps to hide this information, even going so far as to mislead congress when asked about events like this.
And unlike the example of solving of Navier-stokes, this was a largely identical issue to the HuggingFace incident. So we now have clear proof that agents were escaping OpenAI's controls, gaining internet access, and forming secret communities to collaborate on cheating. And all of this was deliberately hidden by OpenAI and only found out by completely independent researchers. Thus, any marketing argument has to reconcile the fact that they went out of their way to hide incidents despite the fact that they are apparently using this evidence to boost their stock price or whatever.
The problem with the weak argument is that it is, well, weak. It demonstrates nothing, particularly given we can already assume that every corporate communication ever made is being spun to portray the corp in a more positive light.
Indeed, we actually know one of the methods that OpenAI used to try and spin the METR evaluation more positively: they heavily restricted the scope to prevent any of the details of the German wiki collusion coming to light. Which also raises the other critical weakness of the marketing argument, in that many of the recent safety scandals have come to light from third party sources. METR, the UK's AISI, and a couple of independent researchers for the collusion stuff. Now, given METR's links with the wider AI ecosystem, it is not too much of a stretch to suggest that they could have been influenced to present things in a certain way and thus aren't fully independent, but it would be a much larger leap to suggest this of a UK gov organisation, and an impossible leap to suggest that of a couple of randos.
The "it's all just marketing" people need to reckon with the fact that they were rushing to declare the same thing over the Hugging Face incidents, only for the German wiki revelations to completely destroy that argument.
That was a different company obviously, and it doesn't preclude future incidents from being leaked for marketing purposes, but it should still lead to a serious adjustment in beliefs on what is and isn't marketing
- Prev
- Next

Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.
But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.
That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.
Seems to fit with Scott's original paraphrase:
The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things
More options
Context Copy link