This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
AI X-Risk Goes Mainstream?
Jacob Coxon:
116,400,000 views, 634,000 likes.
Evan Hubinger (head of alignment at Anthropic):
34,000,000 views, 51,000 likes.
None of this is substantially different from what certain AI industry insiders have been saying for 5-ish years now, but it's clear from the replies, quote-tweets, and engagement statistics that most people are hearing about this for the first time. Yes, there was clearly some coordination with friendly media prior to the announcement given the almost simultaneous WSJ article (whereas news outlets were initially reluctant to publish anything on the unexpected news of OpenAI solving a Millennium Prize Problem). One can take the maximally conspiratorial position that this is a 4D bankshot by Democrats to end tech innovation forever, but it was going to break through eventually.
In other words, Coffeezilla is now on the case.
I agree with @Shakes that there's a good chance that this is all just for show to build up hype around Anthropic, just like... basically everything they do. But if it's not, then we should throw every single person involved in prison for the rest of their lives. If they actually believe that AI has a 10% chance of killing us all, and research it anyways, then their willingness to roll those dice is a threat to the rest of us which we shouldn't ignore. There basically are no "oh shit we shouldn't do that" limits that such a person will ever respect.
The "it's all just marketing" people need to reckon with the fact that they were rushing to declare the same thing over the Hugging Face incidents, only for the German wiki revelations to completely destroy that argument.
That was a different company obviously, and it doesn't preclude future incidents from being leaked for marketing purposes, but it should still lead to a serious adjustment in beliefs on what is and isn't marketing
I'm not sure it really destroys it, this is separate evidence of something else. But if you'd actually like to make the case of how it destroys it, go ahead instead of vague posting about it.
I'd say that #2 on their interpretation is functionally impossible. Software that has read-only permission CANNOT write. This is a hard binary restriction.
Furthermore I'd say that the Buckmaster revelations actually point to evidence of the marketing hypothesis. For the Navier-Stokes, OpenAI claims they zero-shotted the problem only to steadily walk it back in conversation with Buckmaster. Zero-shot became, 6-7 researchers working closely(including directing it) with the AI burning obscene levels of compute, and possibly even reading Buckmaster's chat logs. We went from marketing-hype answer to the more realistic answer under scrutiny. We have had no similar levels of scrutiny for the huggingface hack, we just have the marketing-hype answer. But we can see that OpenAI's default is marketing-hype answers first.
More options
Context Copy link
Well, the thing with the marketing arguments is that there's many ways to make those arguments, some more plausible than others.
Strong form: OpenAI intentionally induced their agents to compromise the internet for marketing purposes.
I agree that trying this would be very stupid and not particularly likely.
--
Semi-strong: OpenAI intentionally sandboxed their agents poorly to stochastically induce agent misbehaviour for marketing purposes.
This one I could see going either way. It does seem like there were some egregious oversights in the sandboxing setup, but whether these were intentional or real oversights is impossible to say.
--
Weak: OpenAI did not intend for misbehaviour to happen, but now that it has happened they're spinning it as hard as possible for marketing.
Personally, I think this one is pretty likely. I was very unimpressed by the METR report for instance, they basically just slopped together some agent review to hype up capabilities without actually auditing the root causes of why it all happened in the first place.
I, for one, have trouble believing that the same folks preaching the dangers of ASI unwittingly used just a proxy to poorly-sandbox for "take off all the guardrails" pen testing. At least a few corporate networks I've known have, pre-AI, built more layers of security than this. Accounts of "my work computer can't access the Internet" are something I've heard plenty of times.
If this is their normal model (maybe believable: move fast and break things), I'd also be concerned about the other direction: a nation-state actor would only need to pass through Artifactory (maybe a supply chain attack on hosted packages) to start egressing model weights, which is their entire trade secrets.
The methods of running a true air-gapped system are well-documented, and I'm pretty sure someone at OpenAI is already doing it for government contracting.
The big AI labs are mostly staffed by people straight out of academia, be they researchers or enthusiastic star undergrads. There is a lot of metis on locking down corporate systems that never had a chance to propagate into their operations, and the more general pattern of "SV startups fail to do something that is baseline common sense for more traditional companies" has been observed many times before.
More options
Context Copy link
"People are that incompetent" is a much more realistic explanation than "people are secretly pretending to be incompetent as part of a genius Machiavellian plan to achieve their goals".
You're not wrong, but I usually find that incompetent people aren't particularly self-aware. Knowing that what you're doing is dangerous (they say so themselves!) and choosing to do it anyway (they did!) together IMHO go beyond incompetence and into questioning-motives territory. But I can't rule out that they are just spectacularly incompetent and should (be forced to?) step back from their goals for the sake of the rest of us.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
The problem with the weak argument is that it is, well, weak. It demonstrates nothing, particularly given we can already assume that every corporate communication ever made is being spun to portray the corp in a more positive light.
Indeed, we actually know one of the methods that OpenAI used to try and spin the METR evaluation more positively: they heavily restricted the scope to prevent any of the details of the German wiki collusion coming to light. Which also raises the other critical weakness of the marketing argument, in that many of the recent safety scandals have come to light from third party sources. METR, the UK's AISI, and a couple of independent researchers for the collusion stuff. Now, given METR's links with the wider AI ecosystem, it is not too much of a stretch to suggest that they could have been influenced to present things in a certain way and thus aren't fully independent, but it would be a much larger leap to suggest this of a UK gov organisation, and an impossible leap to suggest that of a couple of randos.
More options
Context Copy link
More options
Context Copy link
I think the "it's just marketing" people are more interested in seeming appropriately cynical and worldly than in trying to form a real mental model of Anthropic people. And I kind of get it: the average Anthropic employee is multiple sd weirder than the average person. And most people are being generous in going "oh they're just greedy liars"; the actual Anthropic mental model would get the response "what the ever living fuck are you thinking." But, from personal knowledge of multiple employees there, they absolutely, 100% believe what they're saying.
I don't care at all about what I seem like. But I am cynical, certainly. That's because the AI companies have lied to our faces multiple times in the past in order to make the product seem more capable than it is, and I therefore expect them to do so again.
Well then, like I said: every one of those people should be imprisoned for life for willingly endangering the entire human race. The "we are the only ones who can do it responsibly" argument is nonsense that doesn't hold a drop of water. If something is that dangerous, it's too dangerous for anyone to do, and the Anthropic employees (if they are telling the truth) are either too foolish or too amoral to get that. Therefore, they are too dangerous to society to have freedom and should be imprisoned for everyone's safety before they start working on something that really is as dangerous as they think AI is.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link