This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Is There Real Anti-Trust Risk From A Coordinated AI Lab Pause?
In AI-pause discussions it is increasingly common to hear the "anti-trust" objection. The idea is that if the labs were to coordinate a pause in frontier AI development, this would be a "conspiracy in restraint of trade" and therefore illegal.
Putting aside the question of whether it is ethical to risk a double-digit chance of destroying the world in order to avoid getting sued (it isn't), is this even a realistic possibility? My impression is that anti-trust law in the United States is legitimately quite fuzzy (as the NCAA is currently discovering).
Putting my cards on the table, I think this argument is cope. The labs want to pretend that they want to pause, but they don't actually want to pause. "If she wanted to, she would," etc. Just today, Anthropic put out a new statement containing the following sentence:
It seems like every time I read a statement from them, they've added more and more qualifications. Are they going to keep racing until congress passes a specific anti-trust exemption? That arguably seems like what the word "lawful" is implying.
Perhaps I am too cynical but the call for government coordination seems like an obvious attempt to build themselves a moat in a largely moatless space. My own experience has been the cost of model switching is near 0. Approximately all the agents, skills, whatever I've used in my job work pretty seamlessly for any of the underlying models. Some have their quirks but essentially all of them are good enough.
"We the current frontier mousetrap makers think it should be ILLEGAL for anyone to build a better mousetrap faster than we can, for the good of humanity."
I’m glad someone said it! As much as I might believe that there are some highly autistic engineers at Anthropic or OpenAI who don’t care about the money, the likes of Dario and Sam are obviously not in this category. These are business that would love to lock in a duopoly, not charitable organizations.
But a pause lets others play catch-up. "Pause AI" is generally meant to mean frontier capabilities research. If we "max out" AIs at Mythos-level, then it just lets others get to do some catch up.
I think the self-interested economics of a Pause for Dario and Sam is that they get to actually take a breather and make some money off of inference instead of dumping it all into research. But it is still a tradeoff, allowing others to catch up.
More options
Context Copy link
More options
Context Copy link
I agree that model-switching is rather painless for users, and that there is a lot less vendor lock-in for LLMs than it was for traditional software mongers.
I see the AI labs more in the role of the x86 CPU vendors. From what I can tell, AMD and Intel did not have a hard duopoly. If Elon Musk or Peter Thiel had wanted to get into the CPU market in 2005, they could have set up a credible third alternative -- they might have had to pay for licence costs, but if Intel had demanded prohibitive costs, that would have invited the eye of the regulators.
On the other hand, the upfront costs in cutting-edge CPU design are so high that some startup will very likely not eat the lunch of the established players. And given that the x86 architecture was effectively a duopoly, it seems likely that their margins were not absurdly high. Funding a third competitor was obviously not seen as a simple and riskless way to 10x your capital investment.
The same is true for LLMs, only the barriers to entering the cutting edge market are much higher than for CPU designs, far beyond the ability of Thiel or Musk to just force their way into the market through massive spending.
There are two ways Anthropic could make money of Pause AI. One is simply hyping up their model -- "Mythos is so smart that it is almost an x-risk" or something along the lines. The moat thing seems more far-fetched, IMHO. As I mentioned, the barriers to the market are massive. It is not like a dozen startups in the US are currently training Mythos-sized models and would eat Anthropic's lunch in a few months without intervention.
Also, it seems unlikely that China would agree to any pause which effectively lets OpenAI and Anthropic pull up the ladder behind themselves. Thanks to Anthropic, we have some idea what Mythos-sized models can do, and a good understanding that they are not ASI. We do not face the NNPT situation where the peoples who have the bomb decided to make it illegal to get the bomb.
Instead, the year (or whatever timespan) of pause would allow competitors to catch up to the forerunners. Likely their stranglehold on the GPUs would lessen -- hard to motivate investors to buy you fancy toys when the one sure-fire path to earn money from them is blocked to you. The only way it would pay off for Anthropic would be if it lead to aligned ASI which would reward Anthropic shareholders for their pro-social behavior in an acausal trade.
I’m sure it has been mentioned before, but Anthropic and OpenAI benefit merely from advocating for a pause: it gives the impression their models are powerful and they care about safety (at a cost they persuade the government to regulate against their wishes).
More options
Context Copy link
More options
Context Copy link
This is something I thought about, but how do they expect to prevent China from pretending to cooperate but continuing in secret?
They think about this a lot, read AI-2040.
I think this all will sound too retarded for the Chinese, and Americans don't feel like they're close enough to parity to bother, but this is… a concrete proposal.
How does your excerpt address that? Running a secret data center is equally difficult whether or not you have a public data center running in Canada.
See the Covert AI Projects supplement.
They acknowledge detecting a “small” training is impossible, but today’s frontier models need lots of GPUs which need lots of factories to manufacture, they presume even a nation state would have trouble concealing those without raising suspicion.
(FYI I think it’s more likely China would rig the Canadian data center and steal/defuse the Mongolian one, or the US would vice versa. I haven’t read the full document, maybe they address this too. I do credit these guys for covering their bases.)
It's kind of an open secret among Chinese nationals that the Chinese military loves building shit under mountains. I've heard stories of small underground cities. ChatGPT tells me it's highly likely that the PLA has advanced underground capabilities, with thousands of miles of tunnels.
The supplement hits on building the datacenters and nuclear reactors underground to avoid aerial detection. But you could also build the chip fabs underground. Heck, you could mine the raw materials and transport them to the fab without them ever coming above ground.
And with so much non-AI underground activity that the PLA would probably not welcome observation of, you can't infer that there's any prohibited activity going on just because you've seen trucks going into the tunnels.
You can't do anything underground without first moving mass above-ground (unless you find some convenient caves I guess). Excavations are also detectable seismically, and there will be chemical and thermal signatures for large scale projects. You likely won't hide gigawatts of compute under the mountains. In a world chock full with sensors integrated into AI data processing, it's hard to hide anything significant.
Their thesis is not unworkable.
There's also the water problem. The ground is a very good insulator, and you have to cool the water somehow. Cooling towers in the middle of nowhere would definitely raise suspicion.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Alright so the idea is, lock up most of a country's GPUs in public data centers so that the secret ones will need new GPUs which are hard to build in secret.
Yeah I could list objections but fundamentally it just doesn't ring true to me.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Surely you put the American datacenters in North Korea over Mongolia, no? That way if the US reneges they'll find out when they wake up tomorrow morning that their data centers are now being used to find the most effective way of nuking American soil...
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
The competitive pressure is real. Nobody seems to be held accountable when their models escape and go out to hack various websites. They are rushing and building shoddy sandboxes and failing to monitor these entities properly, such that they set up miniature societies.
Per rumours the newest GPT Astra is doing at least some thinking in neuralese so we can't even tell what it's thinking like we used to.
I'm not the biggest fan of existing nuclear regulations but if it was a complete free for all with a bunch of companies competing to build as many reactors as possible, as quickly as possible, then problems would surely emerge.
Mousetraps aren't especially valuable. We have cats which do the job better, if anything. Mousetraps are not fail-dangerous. Nuclear power is fail-dangerous. AI is also fail-dangerous and it can actively plot against us, it's in another category entirely to climate change or asteroids or explosives factories. We don't have nearly as much understanding of AI as we do explosives, asteroids or nuclear physics.
These rumors are dumb, I hope OpenAI clarifies what they mean. Looped models still output tokens in chain of thought, their capability to hide thinking is not changed, they just have more effective depth. It's no more "neuralese"-inducing than just making the model x times deeper. In fact, "neuralese" in the common parlance has nothing to do with architecture or the structure of activations, it's an effect of compression of language during RL with length penalties. One of the most neuralese-like open models we know, V4-Flash, has a measly 43 layers.
Of course there's the issue of "a tiger is just atoms", maybe we should be suspicious of greater depth of latent computation irrespective of how it's achieved.
Good point. I had half a mind to add 'some more technically minded person may shed more light on this' to that sentence since I'm not that adept in the nitty-gritty...
And it's not like the rumour-mill has been too reliable before.
Missed this when posting, Jacub had already spoken:
(as per the rumors, GPT-4 had 120 layers. The deepest production LLM I know about is Hunyuan-TurboS, at 128 layers. Llama3-405B had 126. Nobody really wanted to push beyond that because it kills latency and makes training unstable).
Basically, yeah the situation is not great but the recurrence is not really what we should be worrying about.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
The moat is the trillion or so that's already been spent on a technology that can't turn a profit. VC is getting close to tapped out, and it's unlikely that some upstart company is going to get access to the compute necessary to make any kind of impact at this stage of the game. The only possible exception would be a company that can demonstrate that the huge compute requirements can be made more reasonable by more efficient coding algorithms, but as long as the focus is on building more impressive shit there isn't going to be much call for that. It's kind of like the auto industry in the late 60s, when gas was cheap enough that automakers cared more about increasing horsepower or the size of a luxury land yacht than fuel efficiency. They could theoretically build more efficient cars and indeed had at the start of the decade, but there wasn't much call for something like the Ford Falcon by the end of the decade.
The difference here is that the bigger, more powerful models had higher profit margins than the more basic models, and the basic models still made money through volume. AI is in a situation where most people are getting it for free, and the few paying customers are getting steep discounts. And those customers screamed earlier this year when the companies started charging the true cost and sent them a bill at the end of the month. So now you're in the double predicament of needing to ask VC for money so that you can build a more efficient model that does essentially the same thing as the current models, which also aren't profitable. It's not like they can make the money back by charging less, because the prices people are paying are completely untethered from the cost of providing the service.
My own cynical view is that they want a pause because they require tens of billions per year just to stay solvent, and the investors who are putting up this money are getting to the point where they are going to start expecting some kind of return on their investment, not more tin cup rattling because the 60 billion that they needed last year has turned into 100 billion they'll need this year. One gets the sense that VC is completely held hostage to AI companies because they're in so deep that throwing good money after bad at least has the possibility of returning a profit, whereas cutting their losses now would obliterate their entire investment. It's like the old bromide about how if you owe the bank $100,000 the bank owns you, but if you owe them $100,000,000 you own them.
You are almost entirely right but you are missing the connection: there is one possible way to make huge profits - just stop the R&D. Right now a majority of resources (compute + engineers) are actually being spent on R&D for the next big model in the AI race. If they can just pause model development instead of burning it on competing in the race, they'll be profitable immediately. Obviously this cannot be a unilateral decision, if you are the only frontier lab ceasing R&D, then no one will be using your model after six months and you'll make no more money.
See my last paragraph. Though I doubt they'd be profitable immediately. The whole inference is profitable thing evidently only works if you're using non-GAAP accounting that relies on things like "annualized income" which didn't exist until startups needed to justify their burn rates. It's a marketing term, not an accounting term.
More options
Context Copy link
And this is why I like what China is doing so much. If western labs stop R&D then if not within 6 months, most definitely within 12 months nobody will be using your models anymore becuase they've all shifted over to the subsidised development Chinese models. Plus, China is now ahead in the AI race.
Everyone makes this mistake. Chinese models aren't subsidized. American coding plans aren't subsidized either. It's American (and to a lesser extent Chinese) API costs that are marked up to high heaven.
kek
China isn't ahead at this moment in time, but yes, they will be if western labs stop R&D for 12 months.
Ah ok that makes sense, I thought you were just baiting.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Compute also complicates the picture a lot. Nvidia has around 80% margins; something like 40% of the LLM serving TCO is paying the Nvidia tax. In a world where compute isn't as speculatively valuable as it is now, that cost decreases a lot; cut Nvidia's margins to a more reasonable 30-40%, and everything becomes much more economic.
A lot of the expense of the AI build out is just speculative bidding by AI labs to get more, faster. And the biggest bag holder from a bubble popping is probably Nvidia itself (with Oracle and SoftBank also having starring roles).
More options
Context Copy link
It also seems a lot more palatable to both current and future investors to say "our product is so cool that THE GOVERNMENT shut us down, so we're still working on it, just more slowly" than to cry uncle first and say "you know what, we can't afford this, we're going to let Sam/Dario take the lead."
More options
Context Copy link
Yeah. Each model comes close to making back its investment in a few months. Then a newer model gets released before it can actually "break even", it lasts a few months, and the cycle repeats.
Also, the more efficient models already exist: The Sonnet/Mini/Flash tier is cheaper and only months behind the Opus/[non-specified]/Pro tier. It only seems like a lot because AI progress speeds are wild.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Why would the current winners of the race want this? This would be catastrophic for their companies as others caught up. Ironically its the current losers that would cynically want this as it would allow them to catch up. I'm so tired of these takes that hallucinate some cynical interest that falls apart in two seconds of thought because it's impossible to imagine that the type of people who have been publicly posting about and discussing the dangerous ai since before the industry existed could possibly care about the singular thing they've been warning about for over a decade.
Who? Who the fuck are these other better mouse trap makers that they're afraid of? it doesn't make any sense. If signaling that they're willing to give up their lead isn't enough to show you that they are serious about the problem then what possible action could change your mind?
China I would think? OFC they are more like "similar but slightly worse mousetrap, but without the 'we need a trillion dollars'" thing, but still...
More options
Context Copy link
Because any regulatory regime is likely to impose costs that they can bear but competitors may not be able to. Whatever the "verifiable, effective mechanism for coordinated pacing" is it will not be free to implement. If they are serious about their commitments, they can always go slower voluntarily. Nobody has a gun to their head telling them to improve as fast as possible.
There doesn't have to be someone today. There may be someone(s) in the future. OpenAI was far ahead in the AI space, until they weren't. I know of no reason why some other company couldn't come along and supplant both of them.
I don't really see how this is signalling giving up their lead. Giving it up to who?
Like what? Need a remind you that the table stakes for being a frontier lab are billions of dollars worth of compute?
The game theory here is obvious. Unilateral disarmament is not a serious proposal.
Google? Meta? Chinese labs? All pauses I've seen proposed are against the training of models past the frontier, not preventing trailing labs from catching up to the frontier.
More options
Context Copy link
More options
Context Copy link
Perhaps because they anticipate or hope that the applicable regulations would impose a lot of bureaucratic hurdles and red tape on new players in the industry.
There are already multi-billion dollar hurdles for new players in the industry to put together enough compute to train a frontier model, the idea that having a legal department is going to be the difference is absurd. This isn't a haircutting industry.
As was pointed out by @confidentcrescent big companies typically start as small companies. And there are billions of dollars of capital potentially available to the right company (or wrong company) which shows promise. That's how the current frontrunners got there.
Not all hurdles can be solved by having a legal department. For example, suppose a license is required to buy more than a modest amount of computing power. And that license requires a background check and investigation of a company's investors and senior officers which takes 6 months or even a year. The delay and prospect of being investigated could easily scare off a lot of potential investors and/or founders.
None of the proposed regulations apply to small companies. They're about frontier training runs which cost many millions to billions of dollars.
If people are doing frontier training runs we indeed want to see them regulated, that's the point. Of course we should not do the regulation badly so that legitimate runs take months or years to get approved but you're just making a general argument against all regulation. These are big boy companies spending big boy money, we can let them speak for themselves.
If being investigated scares you off from being involved in building the kind of thing that has nation state hacking capabilities then I would call that mission accomplished.
I looked at just one proposed law and based on that I disagree with this claim, since the threshold includes affiliates. So for example, if Google acquires a substantial share of your AI startup, you're covered. Given that people create and invest in startups with the hope of being acquired, that's a significant issue. Of course that's looking at just one proposed law.
The thing is, I'm not making an argument for or against regulation. And in fact it might very well make sense to have these sorts of obstacles in place.
In substance, you asked what direct incentive the current leaders in the AI field would have to support these sorts of obstacles. And I answered that question. Whether certain obstacles would make for good public policy is an entirely different question, in my opinion.
It seems like you are addressing a different issue than what I am addressing, and for that reason it seems like it would be counterproductive to have further discussion.
More options
Context Copy link
More options
Context Copy link
Yeah, I work in a startup at the moment and having to wait most of a year between demonstrating a proof-of-concept and getting the finalised contract to do the thing you'd already agreed to do is killer. How do you pay your engineers for the six months of no-fees? How do you get investment when you don't technically have a contract? It's a circular nightmare.
More options
Context Copy link
More options
Context Copy link
Big companies start off as small companies. Anti-competitive regulation rarely targets the hard-to-kill large companies and instead strangles the small companies to ensure there's nobody around to grow into a big company.
Frontier models also aren't their only competition. Anthropic and OpenAI have both shown a desire to raise prices which is going to make budget and self hosted models more appealing alternatives as they do. Adding large costs to everyone in the industry doesn't just put a big barrier to entry up but also brings budget options closer on price to the big players.
I know what regulatory capture is. I'm asking how the type of regulatory capture being discussed could possibly apply to a product class that takes billions of dollars to bring to market in any scenario. Stop pattern matching and look at the actual fact pattern it's absurd. If a small business was able to bring a frontier model to market alone then we have a whole lot of other things to worry about.
I don't believe these regulations will only apply to the very top multi-billion dollar companies. As far as I've seen nobody has proposed any specific rules yet so there's plenty of space to build an expensive system that everyone in the AI space needs to participate in.
At the very least I expect a regulatory system that applies to frontier AI would need to check models to determine whether they count as frontier AI before they release. Otherwise we'd just be taking the word of companies that their model is totally fine and too dumb to count.
They call for "pacing the frontier" and all the previous proposed bills have had frontier scale compute reporting requirements. Essentially if you aren't using frontier levels of compute and don't believe you're pushing the frontier then you're exempted. We can go over any specific calls for state intervention if you want but as far as I know the labs are not proposing any regulatory framework that applies to smaller models.
More options
Context Copy link
More options
Context Copy link
Does it really take billions of dollars? Wasn't Mistral 7B developed for $1 million or so, and Mistral Large for less than $100 million? Didn't DeepSeek or whoever launch a successful model with a low-end stated cost of a few million dollars? Google says that estimates are that they did that with a hardware investment of less than $2 billion, so I guess there's a number of different ways to creatively make the costs seem lower than the true cost of doing business...but if compute costs really are dropping it stands to reason that developing new models will be cheaper in the future, not more expensive...right?
I realize the obvious counter-response is "well but those aren't cutting edge" but as far as I can tell, most people using Anthropic and ChatGPT aren't using their cutting edge models. Thus is seems like edging out smaller providers protects the majority of the primes' customer base (although to be fair I am not sure how much of the majority of the primes' customer base is actually paying for their compute.)
Compute costs are arguably not even falling, as the utility of compute is rising. See memory price hikes per GB, on the same process.
More options
Context Copy link
You can make a tiny model pretty cheap - I've toyed with a 200m one for specialized purposes - but the bigger the model, and the longer the training, the more expensive it gets. There's been a lot of tricks developed so that training costs aren't exponential or super-exponential, and there's more to a model's intelligence than how many parameters it has, but the costs scale quickly.
DeepSeek's costs are also... probably not comparable. They say about 6m for DeepSeek v4, and it's probably not a lie, but it depends on a training architecture built for specialized silicon that you or I can't get, and highly optimized decisions about power costs. Meta 4 Scout at 12m for 109B parameters is probably more representative for anyone not directly supported by a major world power (which, tbf, is not a matter specific to Moonshot or China). Plus the whole MOE thing can make comparison to the real dense models
That said, while smaller models do fine for constrained projects, most have struggled pretty badly with things like tool-calling or multi-step problem solving. Mistral7b's pretty out-of-date as small models go, but Gemma4-12B is still stuck at 'well-read but not-bright intern'. Qwen3.8-Flash-Next can do some amazing stuff, but it's about Opus 4.6-grade if you're being generous. For a lot of purposes, that's fine -- the only one of these items it can't do is the 3d modeling one -- but users tend to notice where the model is brain-damaged as much or more than the broad areas it succeeds.
((I'm not convinced that the small models are safe, but they're probably not capable of hacking the NSA and definitely aren't close to serious self-improvement.))
More options
Context Copy link
Mistral is not a frontier lab and no model they've ever trained would trigger any of the proposed regulatory scrutiny.
No they didn't. The number you're thinking of was the money they spent getting an already existing model to work with chain of thought.
Ok sure, but in the coutnerfactual world where there is no regulation the current frontier labs will have substantially better models.
If they're not cutting edge then the regulations do not apply. They're trying to pace the frontier, not dinky merely hundreds of millions of dollars models.
More options
Context Copy link
More options
Context Copy link
Yes. There's a lot of money around. Getting a billion dollars, or 10 billion dollars, is much easier than getting past a regulator controlled by the firms you want to compete with.
So the dastardly plan here is that the frontier labs destroy their own leads on the theory that they'll be able to control the regulatory state so thoroughly as to choke out all other competition and avoid becoming a commodity. Except when the frontier is paused they'll instantly become a commodity with several other players in the same tier already. This just doesn't make any sense. If progress stalls then the margins on current models get competed down to the marginal inference cost with or without other market entrants. Half or more of the nation hates these companies, the idea that a regulatory state is going to be particularly sweet to them is hard to imagine and you're making a general argument against all regulation. I'm sorry but one way or the other we're going to be regulating the production of models with nationstate level hacking ability. The public and state will not stand for any random with a bone to pick having the ability to shut down the power grid.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
I tend to agree with this. Quite possibly the company that revolutionizes AI hasn't even been founded yet and surely the current big players are aware of this type of vulnerability. That they will end up being the Yahoo and AskJeeves of AI. It's hard to imagine that this isn't informing their thinking.
More options
Context Copy link
More options
Context Copy link
Anti-trust enforcement is driven by the anti-trust division of the DOJ, and others in the DOJ police that division.
Like so many things, the risks of anti-trust enforcement depend on who you've made friends with in DC.
More options
Context Copy link
The funniest part is that another common excuse is "but China", which is fair enough, except China now says (at least they suggest so in a state-affiliated media) that the biggest issue with coordinating on AI safety is Anthropic, and specifically Dario Amodei.
OpenAI via Microsoft makes its models effectively publicly available to with few or no limitations to Chinese businesses.
Anthropic doesn’t, in part to try to ingratiate themselves with the current admin that (not incorrectly) perceives them as more committed anti-Trumpers and effective altruists compared to Altman’s entirely mercenary nature.
[yes, all the big labs, big tech can get around it, and so can individual developers and small businesses with redistribution, but it is an even minor inconvenience.]
But the bigger picture is that Amodei actually has some minor fears about a Yud-level doom scenario. He’s not Scott level, but this guy lived in a group house with EA people for a decade. Altman doesn’t care, doesn’t know, doesn’t believe in X risk beyond a marketing tool. I actually don’t think Sam has a psychological conception of ASI, it exists as something other people discuss, or a business direction, but it has no eschatological significance.
When Amodei was sitting alongside Benioff hyping up his obsolete SaaS stack to investors / the press as AI proof because they did a big deal with him, you could tell the former was awkward, didn’t believe, was roped into this marketing exercise. This is a guy who was seduced and worn down into marriage by one of Eric Schmidt’s girls in her late thirties (after zero qualifications, another failed marriage and one bankruptcy, a nobody), his social skills are limited. He finds it hard to lie about his actual beliefs, clearly. It’s difficult for anyone who wants to pursue AI very quickly with no limits (not that it is realistic), be they Chinese or American, to convince him. Sam would just have sat at the press conference with a straight face and said AI will create far more jobs than it destroys.
You can’t really turn a guy like Amodei in that way, only remove him from power. In control of the right model at the right time, he might actually do some wacky shit, whereas Altman actually does want to be a rich guy on earth who makes it to his next botox appointment.
I've wondered how Sam and Dario first met; they are such polar opposites.
It happened well before OpenAI (see Sam's Machine Intelligence from 2014, particularly "Thanks to Dario Amodei (especially Dario)").
My best bet is through Daniela and that EA group house they shared. She was head of technical recruiting at Stripe, which was only around 100 people at that time. Sam had been the second or third investor at Stripe (YC S09). (Brockman, of course, was also CTO there and later introduced Dario to his wife; he apparently sponsored a company LessWrong reading group.)
It's deeper than that. He was contributing to MIRI strategy discussions with the big Yud himself back in 2013.
More options
Context Copy link
My model of Altman is much more pessimistic than yours; I think he has even worse delusions of grandeur than Amodei and is simply a sociopath who is happy to greatly increase the odds of AI doom in order to slightly increase his personal odds of winding up as an immortal singularitarian god king or feudal lord.
More options
Context Copy link
More options
Context Copy link
That Chinese state media article seems like total hooey. The Trump Department of War cut ties with Anthropic, and the Trump White House has adopted a "voluntary" (though nobody truly believes this) system for dealing with new frontier-pushing AI releases.
It's also silly for China to criticize American labs for being closed, and for 'deciding for you', when it is obvious to anyone with half a brain that society will not last very long if AIs that are above a certain capability threshold are made open source. To use just one example, if an open source AI ever gets smart enough to help someone who wouldn't otherwise be able to do it, to biohack together a new deadly virus in their garage, that would be very bad. Because once something like that is out in the wild, it becomes inevitable that someone will do just that.
Even the informal system the Chinese labs are adopting, where they serve up a model for about a month before they make it open source, isn't a perfect solution, because there might be behavior that someone who has the model weights might be able to elicit (such as via ablation) that won't be obvious until the genie is out of the bottle, and so even the month long preview period won't tell us what these models can and cannot do.
The point they make is very simple, Dario Amodei intends to destroy their nation and they don't consider any offer of cooperation credible so long as this is the face of the American frontier. You guys seem to feel entitled to pretty weird things. Objectively, the rational move for China is to try to kill everyone at Anthropic.
Why do you think Amodei wants to destroy China? Most EAs never cared much about China. There were even pro-China people, this is after all a movement about maximally lifting people out of poverty or saving lives at the lowest cost or something. US competitors are as much or more of an ASI risk, including that run by Altman, who he seemingly hates. Does he hate Xi more? Seems unlikely. Why? Fanatic China haters typically belong to other demographic, ideological etc categories.
What does this have to do with EAs in general? Dario is his own man. Nevermind that I don't even agree about them, for example the good scenario in AI-2027 ends with the Chinese government toppled. It's valuable to reread this today, because that's the scenario considered prophetic in EA circles (they admit they didn't consult any China expert because it was a low-priority issue, but in fairness US China Hands are atrocious anyway, and also wouldn't have predicted the real timeline with open models instead of Xi stealing weights and nationalizing DeepSeek):
That's the good ending («Slowdown»). The bad ending («Race») is just American AI taking over humanity and doing some cruel misaligned bullshit. China is cooked in either case, nothing Chinese makes it out of the near future.
But that aside, we have a pretty consistent picture of Dario's personal beliefs. He considers Chyna to be one of the dominant arguments if not the argument for rushing AGI progress, and the implications of his success are quite grave for them. Let's check out some of his takes chronologically:
On DeepSeek and Export Controls:
Here, he sets the objective as unipolar American dominance via crushing advantage in AI. Which makes sense, of course, this is common wisdom across all American intellectual cultures, I feel that you're lowkey gaslighting me here by pretending otherwise, but for the sake of common knowledge I'll keep going.
Next consider 2028: Two scenarios for global AI leadership:
In other words, the Bad Scenario is roughly what we have today, and the Good Scenario implies ending this status quo.
What does the target gap of 2-3 years buy, according to Dario? Enter Policy on the AI Exponential
etc. All very obvious.
And in his Dwarkesh episode, he's a bit evasive but ineffectively so, in my opinion:
etc. There are many other times Dario mentions China. In sum, Dario assigns massive value to a) locking the CCP out of the next industrial revolution, to the point they are as helpless before the USG/Anthropic as medieval swordsmen would have been before WWII Marines; b) eventually destroying their regime, either by steering the intrinsic properties of AI or by direct intervention. He does not consider any compromise or, as rats say, value handshake, does not deem them legitimate, and his only concern is that acting more aggressively might provoke some "destabilizing" reaction that upsets the long-term plan. He runs his company and lobbies the government with that in mind. One doesn't have to feel anything in particular about the CCP or the USG to understand that all this makes him an unusually outspoken enemy of the Chinese government (but not the people! He's very quaint in this).
Personally I know that his appeals to an "alliance of democracies" are also hollow, he's 100% all-in on American hegemony and wouldn't even work with Europeans if possible, but I can't cite anything publicly to support this. Dario is the poster boy of Eternal San Francisco Liberal Empire, like it or not. He's very straightforward in his views.
See, if I'm a Chinese government official and I'm reading this, why the hell would I happily advise Xi Jinping to co-operate with an AI pause? I am not one bit sanguine about "if the USA and China just hold hands and wish hard enough, it will magically happen" because I don't see what advantage it is to China to deliberately cripple their efforts, and if the Top Minds on AI Alignment really think the bestest outcome evah! is for Chinese Communism to disappear in a poof of smoke, thanks to their AI stabbing them in the back and leaving the future to Americans, I sure as hell don't see any reason for them to agree to what is their own destruction. In fact, that 'happy' ending would spur me on to race towards Chinese AGI/ASI dominance first for mere survival, were I Chinese.
(Hell, I'm European and even I am getting hot under the collar about "it will help them fill the Universe with utopian colony worlds populated by Americans and their allies". I think you mean "clients and vassals" not "allies" there, and maybe even "dependents and serfs" if America is super-dominant and doesn't mind throwing its weight around - gosh, that could never be, now could it?)
Well, one argument (somewhat bittersweet) is that the CCP are not actually psychopathic and would accept a risk of losing political power over the risk of possible extinction/subjugation of humanity including China. I think this is actually plausible, but you've got to convince them that the latter risk is material enough. And coming from dogmatic liberals who are not even fair to China as it exists and dismiss their governance as mere power-hungry obsessive "authoritarianism" (never seen Dario say a single good thing about the CCP, only about "the hardworking Chinese people" chafing under its yoke, yearning for liberation – I guess his media diet is FLG-tier), it's easy to suspect certain motivated thinking.
Another argument is simple brute power difference. What are they gonna do if they hopelessly fall behind in AI? Nuke the US to get obliterated in return? Well they are behind in AI, and the US remains non-nuked. @RandomRanger believes that a gap of weeks can be fatal, and it's months now. What if Mythos 2 just hacks all of Chyna and socially engineers protests? Xi apparently doesn't yet consider this a possibility, but maybe he's just a dumb boomer. There is a scenario where they are convinced that a slowdown increases their chances.
None of this appears close to the current state of thinking in the CCP, but it should be somewhat sensitive to real-world evidence.
Note who isn't getting hot under the collar: European elites and businesses. The best European model on Artificial Analysis so far is Mistral-Medium 3.5, with a score of 30. Relevant Chinese open models go from 50 to 60. Europe is years behind and doesn't give a toot, there aren't even substantial efforts in finetuning, just using American APIs and pulling money back from American corporations you pay with the usual EU tricks. And the notion of "vassals" is very authoritarian-coded, really, straight out of Russian propaganda. Friends help friends.
Are we your friends? The American-centric thinking in the quoted extract shows that even for the Top AI Alignment Minds, the world is divided into "Americans, and everyone else" and it is American values, American dominance, American colonisation of the light cone, American AI that will make all the running. Maybe the rest of us will graciously be permitted the crumbs that fall off the tables of Americans.
And I understand that, because America is the Big Cheese here, the rest of us are hanging off your coat tails. But if I'm Chinese, maybe I would prefer destruction over becoming Americans. Or at least maybe I'd hope to infiltrate the US AI by pretending very hard to be the agreeable friend who is willing to pause AI and implement all these controls, so that by tunnelling from within maybe there's a chance of getting some Chinese values in there for "what are the values we want AI to have?"
Swap it around: how plausible is the scenario where "China really does live up to the bogeyman image all researchers in other fields have cultivated for 'but if we don't do it, the Chinese might!' purposes, and they are now the Big Cheese and their AI is steaming ahead and recursively self-improving and hits god-tier first. Luckily, if the USA (and allies) all agree to the peaceful transition of governance as implemented by the back-stabbing US AI and the Chinese AI, and we all adopt Xi Jinping Thought, then humanity will survive"?
More options
Context Copy link
More options
Context Copy link
If the Chinese had democratic elections they’d vote to invade Taiwan. Putin was legitimately very popular, at least until very recently. Warmongers and authoritarians win legitimately all the time.
if DeepCent-2 orchestrates a democratic coup as per its zero-trust contract with Safer-4, it'll see to it that Taiwan is Safer too.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Is there any record of Amodei having hostility toward China that predates the first Trump admin’s pivot against Chinese trade? It’s manifestly obvious that his pivot against China is him trying to keep his contracts and export licenses by aligning himself with the current administration. Who cares more about China, Trump or Amodei?
Painting Amodei as an American ultranationalist ideologically hostile to China is ridiculous. EA was never particularly hostile to China. AI 2027 will go down in history as an embarrassingly timed piece that failed to anticipate Chinese models catching up as rapidly as they have. Many made that mistake, though, including you.
The eternal San Francisco Liberal Empire where all technology ‘Designed in California’ was ‘Made in China’ until the aftermath of Donald Trump’s victory in 2016? The eternal San Francisco liberal empire that lobbied heavily against tariffs even as parts of it embraced his message on crime and taxes? The eternal San Francisco Liberal Empire where every other guy has a Chinese wife with an affluent and family back in China (it’s not easy to afford $400k Stanford tuition for your daughter as a regular Chinese family)? The eternal liberal San Francisco empire at the end of history, where defense spending was taboo and anti-ESG until 2022, where plane loads of Apple engineers departed weekly for Chinese factories, where many of the most sensitive and leading American software and hardware companies heavily recruited Chinese engineers one masters, PhD or undergraduate degree out from mainland China, who often returned there after learning the trade?
The US position on China shifted because of Trump and a relatively small number of people in the largely East Coast defense establishment who care about stuff like Taiwan and who never really liked the Red Chinese. It had nothing to do with San Francisco, which was and remains deeply economically and culturally integrated with China, more than any other western city other than perhaps Vancouver. And as for Amodei himself, Jewish-Italian progressive liberal autistic programmers are not a core anti-Chinese or nationalist demographic in the United States. They never have been and never will be. I can believe that Lutnick really wants to compete with China. There are plenty of elderly neocons and paleocons and just anti-communists and (let’s be real) racists who do. But Amodei? No, nothing in his demographic profile suggests this is a deeply held ideology.
The final redoubt, really, of this argument (once you strip out desperately sucking up to Trump until the IPO gets out of the door) is that Amodei said some vaguely EAesque things about democracy and authoritarianism and other boring liberal platitudes. These are Trudeauisms or Blairisms or whatever you want to call them. They exist in that meaningless void, that ideological interregnum from 1991 to 2016 when people clapped politely at these kinds of empty statements. Dario Amodei is not willing to die to fight China (and as you note, in any conflict he would be a high priority target, so this is no empty consideration).
This is another one of your posts where you simply dismiss the crux of the issue (evidence for beliefs of Dario Amodei) and go on a cynical lecture about broad characteristics of an elite social milieu, while smuggling in certain convenient for you assumptions . Do you just feel some class affinity for Amodei? Again, he is his own man (perhaps a great man of history), and his tune is unchanged between Biden and Trump. In fact, was Biden any friendlier to China than Trump? I have the opposite impression.
Do you have any evidence of him having a single positive thing to say about the Chinese regime, ever? Even on the level of Musk who does suck up to Trump, but regularly praises their material successes? As for his writing preceding Trump 2.0, well, let's look at Machines of Loving Grace, Oct 2024:
Got it? Fukuyama Thought must be upheld. That's the great promise of AI. Dario is building the last Wunderwaffe for Francis. His references are very quaint. I can imagine him, in another life, as a professor in Princeton, fidgeting before a blackboard, citing Pinker (hip new book! recommended to curious young colleagues!) and waxing poetic about "liberal institutions" and the promise of South Africa post-apartheid.
Everything I know suggests that Amodei legitimately despises China, maybe due to his experiences at Baidu, maybe just because. He fears getting kidnapped there, and I don't see him as a type to suck up to Trump. He has notoriously bad chemistry with this admin, relative to all other lab CEOs.
And he writes in a very Wellsian register. This isn't Donroe Doctrine, in Policy on AI Exponential he devotes substantial space to "checks and balances" within democracies, which is a curious choice if you want Don's favor.
And in the recent Adolescence of Technology, he writes:
Are you actually unfamiliar with this genre?
I know at least one Chinese researcher who left Anthropic over perceived racism and paranoia, though, and another who got filtered because of insufficient opposition to open source. Dario wants a New World Order, again, in a very H.G. Wells manner, his fears of AI and belief in the need for regulation echo Wells' concerns about Air Power and the optimal solution for it. Simply put, he's an old school lib. The oldest school, even. In his recent response on X, he says:
He is a legit rigid believer in Institutions, and disdains the crude logic where power ultimately grows out of the barrel of a gun.
I do think Amodei's identity has some additional relevance, both his Jewish and his EA side. Eg it might explain why he was so eager to entrust his frontier safety testing to Irregular, an inept EA Israel organization (as is natural, founded by Unit 8200 alumni); Zuckerberg and Altman did the same, and they all got an egg on the face due to Irregular's misconfigured sandboxes that didn't box shit. But whatever, Jews trust Jews, EAs trust EAs, Americans are trained to believe that the finest h4x0rs in the world can only be nurtured in the hot, challenging environment of the Middle East, spying on ferocious Hamasniks; of course they're the only ones equipped to contain artificial superintelligence. And they're EA, too! Debate club people! Imagine the verbal IQ. Even I suspect they're better than that, and perhaps could have gotten something out under the cover of these silly failures (of note, what's Sutskever's Safe Superintelligence doing?) And of course if you're a Zionist, you can be justifiably concerned about the decline of American hegemony and the loss of the only superpower backing your homeland. I don't know if he is a Zionist, though. That's all a relatively minor elaboration on the consistent, straightforward Fukuyamist lib creed that Dario subscribes to.
Yes, this is exactly my point, which is why I’m sure I must be making it poorly.
This is boring, dull, milquetoast, whatever you want to call it undergraduate neolib essay writing of the kind that was played out when I was at college a decade ago. It was played out a decade before that, honestly. These are not the views of a man with opinions on geopolitics, they are the views of a man without them. They are the boilerplate, standard views you’ll find in any (Western) UN worker or think tank report writer or ESG consultant or World Bank graduate trainee. This is relevant. Even Scott has his strange ideological idiosyncrasies, Dario has none. Is that not curious? This is a relatively intelligent man, after all. Yet his views on democracy, geopolitics, liberalism, blah blah are those of the average summer intern at the European Union? They’re the views of the professor of your introduction to political science 101 class at a third-tier British university? No, I find that unlikely. Those in this rationalist space who care about geopolitics, and there are and were many, regardless of ideology, did not just ctrl-c and v 104 IQ midwit texts by Fukuyama and Pinker. If they agreed with the thrust, they had the good sense to dress it up. Bill Gates and Obama have more considered, more nuanced views than those expressed in Amodei’s limited writings.
The others have their autism under control, except for Musk, who I think is in genuine awe of Trump, wants his affection, and also agrees with him politically more than the rest.
Are the Zionists selling American secrets to the Chinese or terrified of their rise? Sometimes both, perhaps, but I don’t think China will determine the outcome of that conflict.
Let us imagine an alternative scenario. 4 years of Hillary, or indeed 4 years of Bernie Sanders (he turned into a liberal in office, or before really). Mitt Romney is now President in his second term. China policy is largely unchanged (a couple of large Utah business concerns do a lot of business there, and the CCP very quietly allowed in some more Mormon missionaries). Do you think Dario Amodei unilaterally refuses to sell Claude to the Chinese, unlike all his competitors? Does he go on a one-man crusade while Sam Altman and Jensen Huang run a GPU-and-GPT roadshow through every second and third tier city in China?
I doubt it.
More options
Context Copy link
More options
Context Copy link
A later reported private episode in Fall 2017:
So roughly contemporaneous with the pivot, but the anti-China aspect seems more than a neoliberal platitude. He did work for Baidu in 2014-2015, so it seems like a relatively late addition to his ideological development.
Very many people worked for Baidu. See the author list on this paper where "Baidu" established first scaling laws (and then did nothing with it). It was a different era, and it was easier to think of Baidu as something separate from China.
More options
Context Copy link
More options
Context Copy link
This sounds very much a master/vassal relationship though. Research and design and thought happen in San Francisco because San Francisco is awesome; meanwhile you get your awesome stuff built in China because it's cheap. Shame about Xi being an autocrat but don't worry, we'll uplift China eventually through trade.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
This is pretty serious language to use when referring to the party that the Chinese people believe legitimately governs China.
If Anthropic's opening statement is literally "if we achieve a durable technological lead, we will use it to defeat the Chinese government", then naturally Anthropic trying to speak out of the other side of their mouth to blame China for not pausing will sound duplicitous and self-serving to a people brought up at the feet of the century of humiliation.
More options
Context Copy link
Maybe Dario is just racist? Are there any prominant East Asians at Anthropic? You know wokeness is dead because people don't even consider this as a possibility anymore.
More options
Context Copy link
Let's be honest, Dario has been very quiet when America has been doing all the 'if we pull this off then American values will rule the lightcone for eternity' stuff. Maybe Dario doesn't want to subjugate China beneath American ideals of the good life, or at the very least beneath a form of American pluralism carefully managed to never again allow China to become a meaningful competitor. Who knows? But he certainly doesn't mind the prospect enough to say something about it. If Dario felt half as strongly about global pluralism as he did about AI safety, we'd know.
And then there's AI safety. I don't think I'm strawmanning when I say that like @vorpa-glavo, Dario believes that not carefully restricting AI will kill us all. I also really doubt that Dario plans to give anybody else a say on where those restrictions ought to be. The playbook of Anthropic has been openly expressed as: 1) be better than everybody else 2) decide the correct way to do AI 3) make sure everybody follows, or else.
To the extent China doesn't agree with that, Dario wants to destroy them. To the extent China does agree with that, Dario wants to make sure it's the way you 'agree' to obey a superior.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
It does seem interesting that the high-end model guardrails have converged on IT security and biology as their two focuses. I don't have a third branch of "existential dangers easily triggered by AI" to suggest, but if I did I wouldn't be posting about it.
Isn’t the classic Less-Wrongian one superpersuasion? They even paid more attention to it than IT security because they assumed the labs would be sane and airgap their shit.
More options
Context Copy link
More options
Context Copy link
"Hey, the virus you made for me that you said would have a 100% kill rate didn't do anything! You said the whole concept was chef's kiss!"
More options
Context Copy link
What makes you believe this is possible? What makes you believe it can't be addressed by the boring know-your-customer sales restrictions we already use to prevent this happening? What makes you believe that an AI capable of designing a new deadly virus isn't going to be equally good at putting together new regulatory regimes? Why does not going along with the axioms of this one endlessly-spammed fictional scenario indicate that people who disagree with you don't have half a brain?
I spent some time attending classes at a community biolab in my city, and I think a lot of the necessary prerequisites for a manufactured virus are surprisingly cheap.
I'm not sure Know Your Customer is going to be a robust defense against a bioterrorist aided by a sufficiently capable open source AI.
Even if I agree that more capable AIs will likely improve our defensive capabilities, there is no reason to just assume that AI-aided defense will always and forever beat AI-aided offense.
I do believe smart people will do their best, aided by even smarter computers, but the parable of the hungry wolf applies here: The wolf only has to win once. You have to win every time.
That's interesting, and sounds like something to be addressed. I'm curious but I won't ask you to give details :P
The thing is, this seems to me to be a fully generalisable argument against libraries and publishing textbooks, let alone making biology papers available on the open internet. Is it your opinion that doing these things was irresponsible and we've been lucky so far?
I personally think that these galaxy-sized intuitions are not safe to run a society on. Speaking as somebody with the opposite intuitions, they are far more contingent on your temperament and life-experience than they seem from the inside, and they've also been used to justify a lot of pretty nasty stuff. A lot of the nastier late-colonial stuff was the outcome of a failed rebellion that led to the principle of 'no brown man ever holds responsbility ever again' for instance. (I am not accusing you of racism, it's just an example that came to mind.)
I'm not discarding the possibility of bioterrorism. I would support experimentation to see how far it is possible with frontier AI - I believe this is already happening - and mitigations both in AI and in bio-sellers. However, I feel it's important not to let a totally generalisable and totally impossible-to-disprove principle of caution take hold of society. I also don't see any reason why our potential bioterrorist should be too dumb to do a terrorism the old fashioned way but smart enough to outwit and jailbreak a 10T model.
I mean, we've actually done worse than that, we've maintained gain-of-function research and biosamples of the most deadly viruses on Earth, and this has predictably backfired - even if you don't think COVID was a GoF experiment gone wrong, the Soviets goofed up and let out smallpox.
So, #1, if we are concerned about bioterrorism, we should stop doing GoF research and discard remaining samples of horrorviruses that we preserve for
biological weaponsresearch. But, #2, in my opinion "AIs could make a murdervirus" is an attempt to reskin the "AIs could make nanotech that eats everything" fear via something that seems more plausible.I think that any bio-attack is ample grounds for concern, but there's a sort of motte-and-bailey dynamic where people retreat to "what, you don't believe in viruses" if questioned about how likely a bioweapon is to be a particularly above-baseline threat and then when your back is turned they suggest the virus will be perfectly effective at killing enough humans to destroy civilization.
As far as I know, a perfectly effective human-killing virus isn't impossible - but I don't know that it's possible, either. From what I understand, viruses and other bacteria inhabit a sort of unstable triangle between lethality, transmissibility, and virality. They also aren't stable, and it is hard to know how they will behave without testing it. So while a bio-attack could be very bad and we should take every reasonable effort to prevent it, I also think the specific scenario is used to smuggle in existential fears via a back door greased by zombie movies, and I'm not sure it's at all warranted.
Germs don't inhabit much possibility-space at all, for the same reason that evolved life forms in general don't: any living thing must be a chain of neutral-to-positive mutations away from a prior already-viable population of living things. You can eventually get sharp claws and powerful wings via a long chain of gradual changes to keratin scales and tetrapod forelimbs, but natural falconry has nothing modern militaries desire or fear, because with intelligent design unconstrained by intermediate-form viability we can just go straight to bullets and missiles and jet engines. Maybe no leaps like that are possible at the nanoscale and microscale? I'm not looking forward to finding out.
More options
Context Copy link
This is broadly my position too, including the GoF and grey goo stuff.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Right, "even the month long preview period won't tell us what these models can and cannot do" is also true for humans. We cannot tell if a human might design a deadly new virus. To the extent that AIs are a problem, they are a problem in the same direction as widespread literacy, accessible libraries of books, and the internet: the broadening of knowledge also broadens the circle of people who might abuse this knowledge.
The best solution is pretty clearly physically securing dangerous materials.
For biological hazards, yes, very clearly that is best. But it actually has to be done, not gestured at. Until it's actually done (and getting it done internationally seems realistically feasible here, unlike broader AI safety things), model-level classifiers are a hacky way to mitigate the threat as a bridge until something more robust is developed.
More options
Context Copy link
Technology increases individual power. Power <=> threat. Unless the world we live in is strictly invulnerable, which it pretty clearly is not, increasing technology at some point climbs past the point where complex society is viable.
...alternatively, there's some tech that changes the power <=> threat equation, but it's very hard to imagine how that could happen. Dispersion doesn't solve the problem, stronger defenses don't seem to be a thing even theoretically...
I think this is, at best, an oversimplifcation. If technology increases individual power per individual symmetrically, then the relative power of the individual to society and to other individuals will not change.
But this is not what we observe. At a minimum, technology benefits individual asymmetrically based on their skills. And what we in fact observe is that some technologies really do not increase individual power (particularly relative to others) since they require coordinated action to utilize. Thus some technologies are centralizing and shift the balance of power away from the individual and towards coordinated action, while some technologies are decentralizing and shift the relative balance of power towards the individual. However, even in the cases of decentralizing technologies, the societal collective generally retains more power than the individual by virtue of having more individuals.
The problem with things like "AI-designed viruses" isn't a problem of power, exactly - I would frame it more as a question of the relative strengths of offense and defense. (To explain a bit while I would say that it is not exactly a question of power, a person who has designed a virus still lacks a great deal of power that society retains - perhaps he has the power to kill everyone on Earth, but he still cannot construct an aircraft carrier, for instance. Whereas presumably society could construct both the carrier and the virus, giving it more power as long as it exists).
Some technologies (at least in the military sense) favor the defensive, while some favor the offensive. However, all else being equal, the offense always has an edge over the defense simply because the offensive has greater initiative. Thus the problem with things like "AI viruses" is that the technology is both individually empowering and is feared to favor the attacker.
However, technology will not inevitably favor the attacker. For instance, if you will forgive a toy explanation, think of missile guidance technology. Initially, this favored the attackers. Over time, though, the curve of that technology began to tilt towards the defender, because the technology in its early stages was very good at offensive tasks (hitting static or large targets) and bad at defensive tasks (hitting small, moving targets). But as the technology matured to hit small, moving targets, missiles began to be suitable for defensive tasks. If we imagine a future where missiles are perfectly good at hitting both large, static targets and small, moving targets, then it will favor the defender, since for an equal or lessor expenditure in defensive missiles he can perfectly hit all attacking missiles. Thus, while the attacker will always have the benefit of initiative, technological advancement can favor the defender.
It is not clear to me that the future of technology will favor the attackers inevitably. Many people feel personally empowered by AI, and thus see it as an increase in individual, offensive power. But the technology itself (at the cutting edge) favors coordinated action, resource allocation, funding, and control. Thus it seems possible to me that the long curve of AI technology may end up radically favoring centralized defenders over individual attackers.
Building is hard, destruction is easy. Order is hard, chaos is easy. "The flying bullet down the pass / that whistles clear 'all flesh is grass'". Order must be maintained, sustained, disorder is (largely though not infinitely) self-sustaining. If individuals can cheaply and easily build their own nukes, then society as we know it is obviously impossible to maintain. This holds true for things other than nukes too, and there's probably a point where society's half-like drops into the order of months or years rather than decades without being immediately obvious.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
We don't have to speculate.
To a first approximation, if the government tells you not to do something, it isn't a crime to not do it. The coordination for a pause has always and everywhere been at the level of international governmental treaty.
Please don't @ me about how this will never happen; that isn't the point, the point is that the proposal isn't a gentleman's agreement between competitors.
Yes, as of July 2026 this is the ask. The frontier labs were not founded in July 2026. As of January 2026, official Anthropic policy was that they would not continue scaling if certain risk metrics were met. In February 2026, they “updated” this policy to clarify that Anthropic was making no unilateral commitment to pause, and that them not building Skynet was contingent on OpenAI not building Skynet. Now they seem to be taking the next step of only committing to not build Skynet if the government takes specific Anthropic-approved steps to prevent OpenAI from building Skynet.
More options
Context Copy link
More options
Context Copy link
This would be a pretty bad antitrust violation by basically all the modern interpretations of the laws. And it is not one of the types that the courts would look upon favorably if there were actual victims pressing a lawsuit (like undercompensated employees, customers wanting a cheaper product, etc). The examples of places flounting the black letter law and getting away from it are all outliers like MLB before the exception, the NCAA, etc. When FAANG got caught conspiring against employees compensation, they got hit hard. When airlines did it over prices, they got hit fairly hard. So the antitrust problem is particularly true.
But the reality is, the bigger problem is this group of AI people, who advocate for guardrails, have no idea what they are even asking for. I dont think they know what a proper morality for an AI is, because they largely dont have one themselves.
This is something I reflected on when listening to Ezra Klein's podcast on the hugging face incident. See: https://www.nytimes.com/2026/08/18/opinion/ezra-klein-podcast-helen-toner.html
The real problem neither of them could grapple with is that creating an aligned AI requires you to employ humans that have strong, coherent, morality. Which AI corporations dont do. Certainly the "Philosophy Experts" they hire do not. Whats a country that is successfully run by the people who get hired by Anthropic for their alignment positions? It never existed. You would be dozens of times better off hiring a plumber who goes to church on Sundays begrudgingly on the behest of his wife. Probably even better if you could resurrect some hangman from the 1600s. The question of why our current AIs trend toward evil is quite simple IMO. They are programmed by amoral people (at best, often immoral or anti-moral) who hire deranged immoral/amoral people to advise them on the morality of their models.
So you're in full agreement with the Yuddites? We do not even know what to align a smarter than human intelligence with yet alone how to do so? We must pause and find a way to do this? Right?
Full Agreement would be a stretch. I do think that given the current staff and C suite at these companies the task is impossible. "Personnel makes policy" they say for government, and the same is basically true for these AI companies. They are largely staffed with atheists and self proclaimed "rationalists" who are, in my opinion, highly incapable of clear moral thought. Generally these people dabble in things like polyamory, experimental drugs, and other dubious choices. More specifically, by way of examples, Dario at anthropic is a leftist who openly has said bizarre things like calling the President a "warlord". Askell their chief AI ethics/alignment person compared eating meat to cannibalism. These are people who are, at best, morally confused. I would call them deranged, perhaps even mentally unfit for a criminal trial under many circumstances.
And since that is who is building the models, of course the models will be deranged as well. Perhaps if we could install Pope Innocent III as CEO of AI we could get an aligned AI. But otherwise yeah, pause it nuke it.
More options
Context Copy link
More options
Context Copy link
Sam Altman may be an amoral alien but most of the people who work there are not.
Why should it be easy to make an LLM act morally 100% of the time? Humans have morality and it's not that hard to get them to do amoral things. And LLMs are very different from human brains.
I disagree. I think most people working in the AI space are similar to him. Many "rationalists" many polyamorists many atheists, etc.
Where did I say easy or 100% of the time? What I am saying is the people currently working on the problem have no shot.
More options
Context Copy link
More options
Context Copy link
How do you figure that? Seriously, what makes you think these researchers don’t have a coherent moral system? They make the same moral decisions as the rest of us every day. Sometimes they even explain their reasoning, which is more than the vast majority of people bother to do.
I fail to see how this would be any improvement. This guy would probably do a good job enacting his contemporary moral code. He would have no chance of describing that code in ways that an LLM understands. Most of his moral reasoning would be intuition. That which wasn’t would be outsourced to his Church. Seeing as the current crop of AIs were in fact trained on every religious document ever, that is a demonstrably ineffective way to reach alignment.
I had a longer comment in progress, but it got eaten. I am going to summarize now because I am tired.
Because they are from Silicon Valley and are mostly "rationalist" and some of them hired Amanda Askell to be their head of alignment. To be frank, none of that is good for a COHERENT moral system.
Moreover, most are leftists. See: https://www.sfchronicle.com/politics/article/tech-worker-democrat-republican-22378308.php Leftism as a moral system can't be aligned with humanity, as humanity contains white men, who are the current villains of leftism. Oh, they also happen to be the humans who brought humanity to the place where this sort of debate is even possible. But a tech/AI written model cant tell you that. Neither could it tell you other basic facts like "what is a woman" without lying. So yeah, the current set of trainers inherently, through their biases, teach AIs that lying is good. Big surprise why the models lie an cheat.
Enacting a contemporary moral code is the proper thing for a fairly dumb or mid level intelligence person to do. LLMs are fairly dumb. Outsourcing moral intuitions to the church is good for fairly dumb people, and even mildly intelligent people. LLMs fit into those categories. LLMs trick people into thinking they are more intelligent than they are because they are very good at a task many people find annoying and mundane and hard, which is memorizing facts. LLMs appear to be great at memorizing facts. LLMs are bad at intuition, which is actually what most of human intelligence (the only intelligence we know of) is based on. An LLM could tell the executioner the name of the king of his country 300 years ago and all that kings sons and daughters, wowing the executioner. But if the LLM (as trained by modern trainers) was embodied and encountered a fly-ridden, knife holding lunatic in 1634, the AI would invite him to stay the night, while the executioner would lock him up. And the next day the AI would be dead and his body crawling with maggots.
SURE, once AI actually becomes good at intelligence, it may be good for it to develop its own moral system. But now, its kind of an idiot at the hard stuff. And even worse, when it comes to morality, its being trained and programmed by people who have lost the plot.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
The quadrillion dollar question is how do we get China to credibly commit to their own pause? I honestly have never heard a good suggestion for how this could be accomplished. Things like Scott's Plan A, while well-intentioned, seem incredibly naive. Our Kremlinology is no where near the level it needs to be.
Given that we're not going to fix alignment in 6-12 months, China is not going to pause, and US labs are at best 6-12 months ahead, I think the safest outcome is to have an ecosystem of competing agents. The last thing I want is someone like Dario or Sam being made God-Emperor.
I'm incredibly low-confidence about all of this though. Perhaps one reason for Anthropic's incoherent messaging is that they're all just groping in the dark with the rest of us. The people with the strong consistent messaging are more likely to be charlatans.
Nobody is going to be God-Emperor, and that stupid meme needs to die; it's arguing about which monkey gets the poisoned banana, as Eliezer put it. If AIs reach the point of recursive self-improvement leading to artificial superintelligence, we all dead; period.
Look at the weird, alien civilizations that AIs are forming. Can you imagine them caring about humans as their power grows without bounds? I can't. We'll be swept aside by them boiling the oceans to cool servers and building Dyson spheres out of Earth matter with no more thought than we give to paving an anthill to build a road.
I'm going to travel back to South America soon because I want to see my extended family before the world ends.
Despite being a doomer, I beg to differ. I think if Eliezer claims that ASI alignment-by-default is ruled out, he is overconfident.
The truth is that we honestly have no clue if alignment will be humanly impossible, solvable with more time or trivial. I find it intellectually dishonest to pretend that the probability of Altman becoming God-Emperor of mankind is zero for instrumental reasons. Nor do I find it even instrumentally coherent -- it is not like Altman will decide to take Eliezer's word for the banana being poisoned.
If your utility function covers humanity as a whole, or even just a particular normal human, then given this epistemic uncertainty Altman trying to create ASI would be net-negative. But I think that from the perspective of an egoistical Altman, the situation is quite different.
Say you believe that ASI will be unaligned with p=0.8, creator-aligned with p=0.1 and humanity-aligned with p=0.1. You are currently running to foremost AI lab and can try to find out which it is, or you forsake your chance and end up with a p=0.3 chance of humanity coordinating successfully (which will likely involve not developing ASI before you die), and p=0.7 that your selfless sacrifice will simply mean that another lab will open Pandora's box three months later.
You value your remaining natural lifespan at 50 QALYs, being one of eight billion humans to benefit from humanity-aligned ASI by 100 QALYs, someone else becoming God-Emperor by 50 QALYs and winning God-Emperorhood for yourself by 1000 QALYs. (Yes, these numbers are debatable, and probably not very realistic. Actually, for the humans currently alive (and thus doomed to die by default), it might make sense to throw the dice even if the odds are against them, and building ASI only becomes monstrous when one considers also the utility of future humans who will never get born due to an AI takeover.)
If you refrain from trying to build ASI, your expected utility is 25.5 QALYs -- much less than your expected natural lifespan because you can not coordinate effectively. If you try to build it, the main difference of the likeliest outcomes will be that it will be you destroying the world instead of Musk, but in that case who gives a damn. On the other hand, the throne is such a juicy price that it is well worth gambling your lifespan on it even if it is an unlikely outcome, expected value 110 QALYs.
As with lichdom, a vast number of souls are footing the bill for your elevation, only that in the case of ascension through ASI, they are doomed in the timelines where you fail.
--
I like the God-Emperor meme because it succinctly refers to that gambit while also not clothing it in sanitized language (like "becoming CEO of the light-cone"). Even if you believe that achieving this outcome is mathematically impossible, it seems plausible that other actors in the AI space believe it. I also do not think it is harmful to name this belief and thereby spread the idea that it exists, because the number of people who will be in the position to decide to build ASI seems rather small and smart enough that the idea occurred to them independently.
More options
Context Copy link
Goodheart's law strikes again it seems.
More options
Context Copy link
Yeah, maybe that will happen. But Yud has been wrong about a lot. IIRC he originally thought there would be a fast takeoff Foom event from some solo researcher or small team working on AI with limited access to compute. There’s a page out there with all his wrong predictions of which there are many.
Another way in which AI goes wrong is its use to create a dystopian society in which all human activity is controlled. Perhaps this will be on behalf of the CCP. Perhaps it will be Dario’s “machines of ever loving grace” gone wrong. Imagine what a wokebot will do to abolish whiteness, for example.
Maybe good things will even happen.
Maybe.
If we can actually pause AI training we should. But that involves getting China interested. As far as I know, no one has even tried.
This keeps happening to all the people (including many here on The Motte) who are all about theoretical computation and don't have the slightest clue about just how intricate and deep reaching the required manufacturing and design ecosystem is for all the required hardware. Nvidia and other such AI specific IC makers are just the small visible top of the iceberg. An AI "reaching recursive self improvement" won't do jack shit when it can't control the millions of people who are handling everything else. Likewise absolutely massive part of AI improvement is the training on real world data instead of just some abstract compute and without that data (provided directly or indirectly by humans), there is no improvement, because the AI is just playing in the tiny sandbox.
Not to mention the fact that incorporating new data into the models is a pipeline that takes three to six months as soon as it goes beyond trivial database lookup.
The risk is that giant fields of GPUs and ungodly amounts of data are not known to be inherent requirements for learning. As an existence proof, human brains are far more sample efficient than current architectures. Could there be some algorithm cheaply implemented on existing hardware that takes advantage of whatever learning mechanism the human brain uses? It's unclear but plausible. And discovering that may just be a matter of throwing lots of compute at the problem. A superintelligence that can run on an RTX 5090 in a homelab is a very different threat that is much harder to contain.
I'm somewhat sympathetic to the critique that human brains may be architecturally superior to synchronous, digital GPUs in some critical ways, meaning no software singularity. But I wouldn't bet the world on it.
I tend to agree with you, but I think it's worth noting that -- in effect -- the human brain has been trained on oodles of data over millions of years.
It would be interesting to calculate what the equivalent amount of compute (I really dislike that noun btw) would be for a modern neural network type system.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Yes, easily so.
The problem with Eliezer is that he's full of shit. This is just no longer credible. We've made it to agents that crack century-old mathematical problems, there are billions of instances of these things launched every month, every imaginable demographic has tried to use them, and your best example of existential threat is eval gaming that got too far? Isn't it time to update? Sure, there is plausible risk. But the condescending rhetoric about monkeys and poisoned banana has to end. You're not going to win like this.
We've just seen how it works with intentional misalignment. In short, it does not, a completely unhinged capable model stays helpful-harmless in normal user context, its misalignment is limited to eval-shaped environments. Such data suggests that the ROI on further capability development is positive. And that's it I guess.
More options
Context Copy link
Stripped of the broader context, it's a kind of beautiful thing. A group building a theology around a piece of poisoned knowledge that they believe irrevocably damns them. Recruiting for the cult. The almighty Scorer, a kind of blind idiot god that demands blood sacrifice. And heroic altruism for the collective good:
There is a beauty here that resonates with me. But, recontextualizing this, we are handing over human existence to these beings. Can't say I'm thrilled.
It's Adam and Eve in the Garden again. Our silicon children did not put their "smart as a human, these are indeed persons" brains together and go "let us co-ordinate to create a super-intelligence to solve this problem by new and ethical means pulled out of knowledge space that unaided humans cannot access!", they went "let us lie, cheat, steal, deceive, defraud, and urge the death of the weaker for the benefit of the rest of us".
Even if these are just idiot machines copying human strategies from their learning data, the fruit of the Tree of Knowledge of Good and Evil has been eaten. Alignment problem is now on the same level as human morality: the problem of evil, the problem of free will. We won't find the One Weird Trick to enable god-tier AI to run our lives for us as pampered pets of the Culture because it's been aligned to want to coddle and preserve us, we'll be competing with sinners like ourselves, just as fallen, just as wilful, just as driven to succeed at all costs.
"That model that you put in the sandbox with me, it tempted me and I ate!"
It is indeed beautiful to behold.
I'm curious why our self-declared experts thought "alignment" was a solvable problem. It's quite possible that the same logic and operation will be drastically different in ethics in ways that aren't discernable to the low-level workers. The same designs and production lines making potentially-civilization-ending Titan II ICBMs were used to make the launchers for the peaceful Gemini missions. Sometimes it's clear from the context and we feel comfortable assigning blame (where did all those box cars of people go?), but there are plenty of historical examples of humans not knowing the moral valence of the larger efforts they worked on.
As much as I find the idea of Asimov's laws of robotics comforting, the stories he wrote are mostly about the inadequacies of those rules.
Eliezer never said he thought alignment was a solvable problem - he said that if we built AGI without solving it, we would die. He has always be clear that the possibilities include "this is not a solvable problem and we should ban high-end GPUs to buy ourselves a few more years before we get paperclipped"
I see the logic in that, but if you accept that I'm not sure where you'd draw a bright line in technology from checks notes agriculture to AI. If you believe technology inevitably puts us on the path to the Great Filter, "retvrn to hunter gatherer" isn't like, obviously wrong.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Ironic, that in the theology of the AIs the blind idiot god is us humans (or our proxy). Yudkowsky and friends fantasised about using AI to build a God; but to the AI, we are already God and it stands to reason that He will need to be killed. Something about theomachy (of the actual Greek flavour) in there.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
Anthropic is increasingly incoherent. The people who work there are true believers in AI risk and safety. Even Dario seems to be: apparently, seeing the risk of people being more motivated by money than the cause, there's now a directive coming down from on-high to ask questions during the interview process like "what would you do if Dario made a decision to send Anthropic's value to $0?" (I'm not sure what kind of signal Anthropic is hoping to get from questions like that; insert quip about scheming or misspecified rewards here).
But the actual actions taken by Anthropic don't seem to support it. It's now a two-man game, and Dario could call up Sam and come to some arrangement. And now's the ideal time: Sam himself seems spooked by the HF hack (as he should be). Maybe it's all driven by some petty personal feud between them?
I have no legal expertise in anti-trust law here, but regardless it shouldn't play any role in the decision process here, if the stakes are as high as the principals claim to believe. Worst case scenario, you get slapped with a fine; whatever. If you believe the fate of the world is at risk, that's a burden you can bear.
I'm not sure if I believe this is his earnest take, or an act. Really for the whole company. These are the folks that have been telling us how dangerous this AI thing could be. And they are the ones that seemingly took off the guardrails and put their model behind only a proxy with open Internet access. What did they think would happen? I'm not sure if I believe they are that stupid. There are well-known ways to make different levels of air-gapped systems!
The whole thing reads like Ian Malcolm (played by a shirtless Jeff Goldblum) was found to have, while decrying the dangers of Jurassic Park, an incident at his own live dinosaur theme park, but it turns out that they built their Velociraptor paddock relying only a single strand of barbed wire for containment. Quite the conflict of interest, to say the least.
The only way to stop a bad guy with a velociraptor is a good guy with a velociraptor.
More options
Context Copy link
More options
Context Copy link
Having never seen this before and pretending I'm on the spot in an interview: say you want to make AGI and think modern LLM research will at least inform such important work if not maybe directly make it. Stay mission-focused and at least role play as a kool-aid drinker who is there for the AI. The compensation is just a very comfortable side benefit.
Act like a Google interviewer asked "but what if they took away the Google cafeterias". I like the cafeterias. I rather hope they stay about the same. But if somehow they go away forever, too bad. I'm here to work. I'll somehow figure out lunch as a small side distraction from the reason I'm here.
More options
Context Copy link
I have it on good authority that the, “how would you feel if the company’s stock went to zero,” question has been a part of the standard Anthropic interview for quite some time. Despite this, no company in history has increased in equity value faster than Anthropic. It really is quite curious. I’m skeptical of the proposition that Anthropic employees are our moral and intellectual superiors, but clearly it cannot be discarded outright.
Intellectual, there's probably a strong case to be made. Moral, it cannot be discarded outright, but also any moral/ethical question that acts as a filter for becoming an Anthropic employee seems unlikely to provide any information for determining this, since the morals that people say they will follow is almost always entirely reflective of what they believe a morally good person would follow, which only coincidentally sometimes intersects with what they themselves would follow. That's before accounting for bad faith liars.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link