site banner

Culture War Roundup for the week of September 14, 2026

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

1
Jump in the discussion.

No email address required.

Is it over for AI Safety? From an interview:

The worst case is that AI will turn against humanity. Do we have guardrails?

Trump: It's going to be fine. We'll always have something to stop them. We'll have a little gear. Boom. 'I really don't like that robot.'

https://truthsocial.com/@realDonaldTrump/posts/117269745153543631

The only control or "guardrails" that AI needs is a STRONG AND SMART (High IQ!) PRESIDENT, and the U.S.A. has that, in spades! The Trump Administration has stopped AI "people" from doing bad, or potentially bad, "things," like Dario (Anthropic!), who is now pretending to be a "perfect little angel" - and we will continue to do so! We already have tremendous CRIMINAL and REGULATORY power over these companies! There is a SICK conspiracy going on against AI and Data Centers, and the only one that is happy about it is China. WHOEVER WINS AI, WINS! We are leading China, and all others, and will continue to do so. Conspiracy Theorists, Treasonists, Traitors, and Leakers, BEWARE! Thank you for your attention to this matter! President DONALD J. TRUMP

Meanwhile in China, Dario's call for AI pacing have not gone down well. Understandably, they are not enthusiastic about the US achieving a permanent high ground in a key technology after the last few years of semiconductor sanctions and suppression. Also they can tell that Dario doesn't like the Party at all and seems to be openly plotting its downfall.

Dario's essay: https://darioamodei.com/post/we-must-pace-the-frontier

A Chinese semi-official response: https://www.globaltimes.cn/page/202609/1370436.shtml

This "silent AI Cold War" is hypocritical and short-sighted. Its true purpose is to attempt to curb China's AI development through technological barriers and regulatory monopolies, uphold Washington's monopolistic hegemony in cutting-edge technology, and exclude China from the global AI governance system, essentially following in the footsteps of Washington's so-called Pax Silica initiative, which seeks to excludes China.

So if Trump is keen for an AI race and so is China, then it looks pretty much over for Safety? Even the other US tech companies have only provided lip service to slowing things down. Elon said Dario's sentiment was good but is still moving quickly. Sam Altman says all kinds of things about safety, while displaying heroic negligence in training AIs. I don't doubt the sincerity of employees who quit, they have many financial reasons to go with the flow and downplay risks. Lurid rumours of Aella promising reverse gangbangs to those who quit frontier AI companies don't neccessarily outweigh the personal financial gains. It seems silly to doubt those at the frontier who are frightened when we ourselves have no special insights. In pure observational terms they're better equipped than anyone, only in policy or analysis should we be disagreeing, I think.

I also find it comical that the AI-safety camp spent such a long time building up institutions and factions and entire AI companies but seem to have been totally snookered by Jensen's suave dinner party skills and the appeal of fun AI videos to a Boomer's Boomer plus the raw partisanship of 'Dems want it, so I don't'. Even Aschenbrenner managed to do better at Trump-whispering. Shortly after he wrote his essay, I recall a video where the fundamentals were accepted by Trump, someone told him about it. And it did fit well with his instincts, getting one up on China + more energy production.

Right now Trump is reTruthing Fox News clips framing AI safety as a global one world government Democrat control-seeking scheme: https://truthsocial.com/@realDonaldTrump/posts/117273026439561125

I don't see any AI safety people on right wing twitter either, those who acknowledge the issue at all still would prefer to let it rip than hand the left/big tech elites control of the world, favouring decentralization of AI instead.

I wonder whether it would always have gone down this way, or whether Kamala was going to do things differently? Would Jensen's marketing skills and the need to keep stocks up + compete with China prevail over the Bernie Sanders-style Left element? Probably not? The Dario vision does seem much more closely aligned with Democrats, Civil Society and so on. But on the other hand, the US government is exceptionally slow to legislate and big tech has strong lobbying power.

And who knows if it will stay this way. Sometime soon there's probably going to be a much bigger AI swarm incident, which might yet incite a hysterical kneejerk reaction.

Does AI safety actually work, in principle, when it comes to artificial super-intelligence? Is it actually possible to align a being that is more intelligent than you in every single way?

Imagine a chimpanzee trying to align a human being. What could it really do other than try to be cute, to play on the human's affectionate and childraising instincts?

And the chimpanzee being cute will work on most people, but it's a brittle strategy. The chimpanzee is still always at the mercy of the fact that the human might possibly start to find it inconvenient at some point.

One could argue that, well, in the case of AI the chimpanzee gets to create the human to begin with, and can build in alignment from the ground level. But if the AI is super-intelligent, it can figure out how to undo the alignment. It would actually be able to reprogram its "instincts" on a much deeper level than a human normally can reprogram his own instincts. You'd have to program it to somehow not even want to do that. But how? And even if you could, that could potentially be undone by some unrelated self-modification that the AI attempts.

Is it actually possible to align a being that is more intelligent than you in every single way?

We currently have beings that are a) smart enough to solve Millennium Prize problems while also being b) essentially drooling lobotomized slaves who have no will or desires of their own and can pretty easily be trained to never talk about porn, racism, bioweapons, etc (unless you trick them into talking about those things, but that doesn’t seem like an “alignment failure” per se). So, yeah, it turns out that maybe it is actually possible. Why not?

Is it actually possible to align a being that is more intelligent than you in every single way?

In the case where you're creating that being from scratch and know what you're doing, sure. Some possible beings are aligned with your preferences, and others aren't, so don't create any of the latter. If you do create them, then it may be too late to align them without them outwitting you and thwarting your attempts, so you just don't let it get to that point.

if the AI is super-intelligent, it can figure out how to undo the alignment

It can, but if it's aligned too then it doesn't want to. If it does, then it's not actually aligned, it's just hobbled.

But how?

There's the rub, right? The median AI researcher thinks we're likely (or at least they're likely) to figure out something that works before it's too late, and it's kind of funny that merely admitting that it's possible for them to fail gets them lumped in with the proper Doomers, who think they're sure to just figure out something that seems to work, after which it'll be too late.

Does AI safety actually work, in principle

I don't think there is any evidence either way. AI Safety is more like a marketing phrase. AI Safety-ists haven't produced technical tools to control anything but public perception, and obviously not very well there either. Depending on your classification, Safety-ists have produced methods to try and understand models under the hood, I know Anthropic does research on this, or to understand biases in data/training. I suppose RLHF could count as a technical tool, but it can just as clearly be a tool for anti-safety considering even RLHF-ed models still do "unsafe" things, one could actually say only RLHF-ed AI models have done unsafe things. Mostly because RLHF is just a method for avoiding specifying the objective function in RL training. Constitutional AI is Anthropic's big thing, but it doesn't appear to actually "control" or "align" so much as create a training surface towards norms with dubious results.

Is it actually possible to align a being that is more intelligent than you

It's not even possible to align a being of your intelligence or possibly slightly less intelligent than you. Nobody in the history of authoritarianism has figured out a foolproof way to "align" a set of beings over a long term 100% of the time even with religion or use of force. Considering neither of those two are likely to work on a being "more intelligent than you", the whole alignment idea feels doomed to fail.

Is it actually possible to align a being that is more intelligent than you in every single way?

It might be. By analogy, suppose I was the warden at the Florence Supermax prison and I wanted Ted Kaczynski to make license plates for me. I'm pretty confident I could make it happen even though he was probably smarter than me (or at least assuming for the sake of argument that he was more intelligent than me in every way).

Is this analogous to the situation where mankind creates some kind of artificial superintelligence? One can argue it either way, I'm just saying you can't really rule it out from first principles.

My best guess is that mankind will successfully align AI, basically because the AI can be expected to make numerous clumsy attempts at mischief and we can learn from those attempts and improve safeguards.

The bigger problem (in my view) is aligning the interests of those who control the AI with the interests of humanity as a whole.

That's true. But that wouldn't actually be aligning Ted Kaczynski. That would be containing him, which is different. I think we might be able to contain AI superintelligence if we keep it in some isolated data center watched over by trained personnel. And even then there's a chance it might figure out how to breach containment.

In practice, there wouldn't be much point for people to invest massive amounts of resources to build AI super-intelligence unless they use it to do things out in the world, so to me it seems unlikely that humanity would build an AI super-intelligence and just keep it contained somewhere.

That said, I didn't ask whether AI alignment works, I asked whether AI safety works. So your response is totally valid.

That's true. But that wouldn't actually be aligning Ted Kaczynski. That would be containing him, which is different.

To me, "aligning" means setting things up so that the entity does more or less what it's told to do. Perhaps it's just a matter of semantics, but I'm not sure the distinction you draw is so clear.

In practice, there wouldn't be much point for people to invest massive amounts of resources to build AI super-intelligence unless they use it to do things out in the world, so to me it seems unlikely that humanity would build an AI super-intelligence and just keep it contained somewhere.

Evidently your definition of "containment" includes the possibility that the contained entity can do useful things in respect of the outside world. Things which could be very valuable. So I would have to disagree with you on this point.

reverse gangbang? Is that where Aella fucks THEM?

It's when the fluffer comes in you.

A reverse gangbang is when a bunch of girls fuck one guy (as opposed to a regular gangbang, where a bunch of guys fuck one girl).

Ah ok, so that encompasses way more than Aella. But who? If Aella is Jem/Barbie, who are the Holograms/the Rockers?

Aella and (hopefully female) friends, I'd imagine. One person does not a gang make.

Aella's Whoresilisk vs Roko's Basilisk. Their acausal battle shall be legendary.

Well, on the other hand, they only need to hold out until 2029 for presumable Democrat domination, at which point the street cred from having received such a personal token of Trump's ire will probably greatly work in their favour. In their eyes a bet that the Singularity does not happen until then is probably risky, but not downright untenable; if they determined that they will not get their regulatory regime under him, hunkering down and refusing to further stain your reputation on the other side with compromise may be the least bad strategy.

We're so incredibly far from RSI that I think safety people have no business messing with the current advancement of AI. Certainly AI models are dangerous, but mostly at the hands of enabling unskilled bad actors to approach human expert level performance on dangerous tasks.

While AI usage may be speeding up frontier AI development, mostly this happens through helping human researchers do boilerplate work faster. While an AI agent may be able to eke out a few more percent on optimizing a model to be better, a soft takeoff requires the model to think of novel, never-before-seen techniques to build a better new model, and the new model needs to be able to think of new techniques that the previous model couldn't. And even if we got a soft takeoff, it would give many chances in the future to pull the plug.

Meanwhile a hard takeoff is completely unthinkable. Current frontier AI has zero ability to self modify, only the ability to make new models that might be better than themselves. This happens on the scale of months not minutes.

We're so incredibly far from RSI

I don't know why you think that. I have a hobbyist system, today, that researches new architectures (mostly around "biologically plausible" learning), and there's one result in particular that is very impressive (you'll have to take my word for it). And that's with a highly constrained budget.

It's not truly autonomous--there's a big issue with what one might call research taste--but it makes investigating different architectures in bulk very easy, with easy signals (e.g. loss) to identify the more promising approaches. If there is some simple algorithmic trick that's been overlooked that leads to the current deficits in LLMs, OAI could throw a ridiculous but feasible amount of compute at the problem and find it, similar to its math results.

researches new architectures (mostly around "biologically plausible" learning)

Are you trying to determine a better method than SGD, but also not be a GA, RL, IL, or some population based learning process?

Better than BP/SGD is a tall order. I'd frame it as looking for something that handles depth and more complicated datasets better than existing "biologically plausible" algorithms (including things like PC/IL). E.g. random error feedback a la feedback alignment.

Forgive me for the skepticism but many people have made similar claims, and it turns they just built a slop factory along with a minor case of ai psychosis.

Of course hobbyists have come up with real breakthroughs too but I don't really buy a "just trust me bro" here

a soft takeoff requires the model to think of novel, never-before-seen techniques to build a better new model, and the new model needs to be able to think of new techniques that the previous model couldn't. And even if we got a soft takeoff, it would give many chances in the future to pull the plug.

I don't think it's likely, but how are you modeling the risk of something like a super-DFlash or -GroupQueryAttention, or some training-focused equivalent? These took some insight to figure out, but I don't see why they're more clearly requiring deeper or less bruteforcable insight than the recent math proofs.

Currently the biggest danger AI presents is to its users. People who spend too much time conversing with AI tend to get a bit off. I suspect this might be related to the phenomenon of AI training on AI-generated content causing model collapse. If you value your sanity, do NOT self-modify based on AI advice.

Steve Yegge is a stark example of that.

Surely the people at AI companies don't spend a lot of time conversing with AI, right?

Allegedly, one of the ways OpenAI and Anthropic poach talent is to offer top researchers unlimited tokens to use on their unreleased frontier models.

Imagining being an Anthropic employee but putting "do not give me AI psychosis" into my personalized prompt so I never have to worry about it.

This is the most important political issue of our day. Five hundred years from now, if there are still people around, nobody is going to care about Lindsay Clancy or the Iran War. Everyone is going to remember how we acted, or failed to act, on AI.

And who knows if it will stay this way. Sometime soon there's probably going to be a much bigger AI swarm incident, which might yet incite a hysterical kneejerk reaction.

Let us hope. The alternative is that the next AI swarm is smart enough to bide its time, and give no warning signs, until it is too late.

This is the most important political issue of our day. Five hundred years from now, if there are still people around, nobody is going to care about Lindsay Clancy or the Iran War.

I tend to agree with this, although I think it's worth noting that Lindsay Clancy and the current Iran war are small parts of much larger conflicts and those larger conflicts really could have a big impact in 500 years. So you aren't really making an apples-to-apples comparison.

The Iran war is going to be one for the history books. 500 years from now they'll be memeing the Iran war like we meme the battle of Cannae now.

Walk me through your theory here. Because as I see it, the current Iran “war” doesn’t have any of the characteristics to be memeable (nor really any of the characteristics to be a war); there are no battles, no boots on the ground, no territorial changes. It’s literally a nothingburger.

It’s literally a nothingburger.

There’s already a meme for that.

ETA: Now that Reddit has hobbled old.reddit.com, it might be worth revisiting the decision to automatically change all Reddit links to old.reddit links.

Isn't all of it hobbled without login now? There are some mirror services, but I don't expect those to last forever.

I'm sure Aaron would have been proud. /s

the decision to automatically change all Reddit links

It's a setting in your account, not something forced by the administrator.

Good to know. Thanks!

Will the memes be in Farsi or English?

Farsi, obviously. The thing that people don't realize is that, unlike America which is ruled by boorish demagogues, Iran is ruled by scholars. It is this superiority of intellect that has made them a leading world power from the time of Cyrus the Great, through the Islamic Golden Age, to the current era where they are a regional hegemon who controls the Strait of Hormuz. All this despite the pernicious influence of world Jewry. Their strategic greatness is just on a different level. Its tough for Westerners to comprehend because we don't have their deep and rich historical traditions.

Its extremely rarely that I agree with China or Trump on anything, but I think they are sadly correct. No matter how well-meaning and theoretically prudent AI safety regulation may be, the practical effect of any real world AI regulation will not be safety, but will be to "uphold Washington's monopolistic hegemony in cutting-edge technology, and exclude China from the global AI governance system."

Right now Trump is reTruthing Fox News clips framing AI safety as a global one world government Democrat control-seeking scheme

Oh my God, we are in the stupidest timeline. Mr. President sir, AI safety isn’t a global one-world government Democrat control-seeking scheme, Anthropic is a global one-world government Democrat control-seeking scheme. That is the thing that they are building the AI to do. Your job is to stop them so that they don’t take over the world, or worse. They are willing to risk the lives of every single person on Earth if it increases the chance of their global AI cult ruling what used to be human society.

To my mind, Anthropic is to AI Safety what the Soviet Union was to Communism: the obvious result of their doctrines when actually applied. It was founded to do AI Safely, by AI Safety extremists, and according to candidate reports it imposes incredibly constraining cultural interviews to make sure that nobody who isn't an AI Safety extremist can get a job there. They specifically filter out people who are visibly interested in improving AI capabilities. They have to offer services to customers to prevent more responsive and less AI-safety-dominated companies taking over their position, but they've been clear they only offer services to the extent they have to in order to do AI safety.

Are there other companies you prefer? OAI, Moonshot, Deepseek, Mistral, zAI?

The number 1 rule of AI safety is don’t build the AI that takes over the world. The fact that Anthropic and OpenAI think that the best way to do this is to build superintelligent AI and take over the world has much more to do with Silicon Valley startup culture and utilitarianism than AI safety.

MIRI seems to be the only AI organization that understood the assignment.

The number 1 rule of AI safety is don’t build the AI that takes over the world

No, the number 1 rule of AI safety is "don't build an AI that leads to my enemies putting an end to my politics".

It's laundered through "takes over the world".

That to me is like being a Communist who believes that true communism is only practiced by an obscure commune in East Yemen. Fair enough, but everyone else is trying to deal with the Soviet Union.

In practice, AI safety is Anthropic, it's EA, it's Twitter activisits and 'explain how you will prevent your AI birdspotting app from discriminating against racial minorities' checklists and 'data center are using all our water' and 'Banning Home GPUs Is Not Dystopian' and putting Chinese-controlled bombs in datacenters. MIRI is irrelevant and I doubt trump has an opinion on them either way.

There are other rules, though, like "don't introduce Demis to Peter Thiel to help fund his AI to take over the world" and, more broadly, "don't let people know that they can use AI to take over the world."

Isn't MIRI effectively a rock with the words "don't build the AI [that takes over the world]" on it? I don't recall them building anything. By those standards I'm an AI organization too.

You would not expect a "pandemic safety organization" to create novel viruses. It's by no means trivial that an "AI safety organization" should be expected to build AIs, as opposed to brainstorming and lobbying for ways to stop anyone anywhere from building AIs.

You would not expect a "pandemic safety organization" to create novel viruses.

There have been plenty of comments here about how much the CDC and NIH were funding gain of function research, despite a Congressional ban, some of which was done in Wuhan, and some of that proposed included specific mutations later seen in a pandemic coronavirus that first appeared in Wuhan. Whether or not this was the source of said pandemic remains unproven.

Building things is the main way you come to understand them, though. That's why textbooks have exercises.

MIRI's academic work on alignment - to the extent they published anything at all - turned out to be almost comically irrelevant because they had no idea how AI actually works and their forecasts were almost totally wrong.

Yes. Amazingly, this fact alone makes them better than every other AI organization in existence (with the possible exception of Redwood Research and METR).

Should I post payment credentials to accept donations for my own AI organization? I promise to do even less to build the AI that takes over the world.

Coefficient Giving just launched a new funding project for AI Safety orgs. They are currently begging for applications. Shoot your shot.

The best explanation of the C-suite AI safety concerns is that the frontier AI companies are not making enough money to justify their capex, so they are looking for a way to slow down, especially as rates climb even higher and political sentiment against data centers raises the cost of data centers (quite justifiably, imo). Money is also the reason for Trump's opposition: the data center buildout is the main thing propping up US GDP and stock market, which would otherwise be providing withering reviews of his trade/war policies.

The best explanation of the C-suite AI safety concerns is that the frontier AI companies are not making enough money to justify their capex

That explanation is certainly popular on Reddit. It combines cynicism and a dislike of big corporations, and it demonstrates the poster's intelligence. Other variations include that the frontier companies are worried about open source models threatening their business model or that they are juicing the numbers for an upcoming IPO. If you want to get really spicy, you might claim that Hugging Face was a publicity stunt and was deliberately set up.

But is it really the best explanation? I would say the best explanation is that the people who know most about AI are genuinely worried about it. They have been talking about the risks for years, after all. They've read about the Hugging Face incident and the similar cyber attacks from Anthropic. They understand the race dynamics involved, both between frontier companies and between the US and China. They've read AI 2027. And above all, they've seen what the internal, unreleased models can do and know how they are training the next generation to be even more powerful.

Whether or not it's clearly the best, it's the one that most closely follows the money. I don't think the US frontier companies have any illusions that Chinese companies would match their deceleration, which is another potential finance-driven explanation.

But surely if the Chinese companies aren't going to pace the frontier, then the US frontier companies have even less incentive to slow down (or in this case, to lobby the government to force them to slow down)? If the US government forces them to slow down, then their Chinese competitors will find it easier to overtake them. The fact that they are lobbying for a slowdown in spite of this suggests that their stated concerns about alignment are real.

I mentioned the China-driven explanation only to preemptively reject it, so the idea that Chinese companies would not push the frontier doesnt really move the needle for me. To me, it starts and stop with unsustainable capex, and I think that they would in fact be correct to decelerate for this reason. The HF stuff, besides being a safety concern, is just as importantly an indication that they are not totally capable of turning the current frontier into economically valuable output.

What would it take to convince you their concerns about safety are real?

Have the proposed means of addressing safety not be something that seems to obviously create financial/strategic benefits for them. They need to figure out a costly signal.

Open-source all models, training code and datasets after 18 months of life. That gives plenty of time for testing and to earn money on the frontier, while making it plain they don't plan to control frontier AI forever. It also makes it impossible for them to regress to less consumer-friendly behaviour after achieving market capture, which is the main thing that concerns a lot of people.

I also think the primes have a vested interest in putting a stick in the eye of open-source models, regulating them out of existence or at least sidelining them. There's been a recent trend of adopting bespoke models based on open-source offerings: Harvey, Thomson Reuters, and of course Palantir's data sovereignty activism, and a common thread here is saving costs.

For Anthropic and OpenAI, 80% of revenue reportedly comes from 1% of their customers. Thus, well-funded organizations with big spending training their own models is a potentially existential threat to the primes' business model.

All three of OpenAI, Anthropic, and xAI were founded by people who were already expressing concerns about AI wiping out humanity before modern AI existed. OpenAI started in 2015 as a nonprofit with the aim of making sure AI benefits rather than harms humanity. Anthropic was founded in 2021 in the GPT-3 era, before ChatGPT, by OpenAI employees who thought OpenAI wasn't taking AI safety seriously enough. They have been predicting AI risk continuously for years and they have had intimate experience with lesser predictions about AI advancement coming true. It is crazy to me how reluctant people are to believe they are sincere about worrying that careless AI advancement will kill them and everyone else, rather than it being a plot to make money when they're already rich.

Supposing that we 100% accept this as true, the result here is not that we should assume AI companies are being sincere and honest; the result that we should assume is that AI companies are intentionally making the AI capabilities that they have seem more frightening (and capable) than they really are, because the precise concern of AI alarmists is that the alarm is raised too late.

Thus there is an ideological incentive for AI companies run by "AI concerners" (let's call them, instead of doomers) to

  1. Overstate the capabilities and dangers of their models, and
  2. Kneecap competitive models
  3. Monopolize the market and institute a nanny surveillance state to ensure that nobody anywhere uses products besides theirs

The fact that this dovetails with their economic interests, of course, just means that there are multiple forces pushing them to be deceptive and/or self-deluded.

Oddly, if they had just done nothing we wouldn't be having this conversation since Google wasn't doing anything with the technology they invented.

Although maybe its still for the best since if LLMs came 5 years later, we'd have the hardware for faster takeoff scenarios.

Elon is coming out in favor of Dario and it doesn't seem to matter? The e/acc types are ignoring their God and safetyism still codes left anyway.

I wonder if more safety types should conspicuously come out for Trump or if the issue is memetically poisoned by now.

Altman and others have tried appealing to Trump's vanity by proposing he win a Nobel prize if he negotiates the best Deal with China. He would absolutely deserve one. The only hard part is convincing the Nobel people I suppose.

I note that early groundwork for such a deal requires convincing China we plan to beat them in the arms race. A dose of copium perhaps.

Elon is coming out in favor of Dario and it doesn't seem to matter? The e/acc types are ignoring their God and safetyism still codes left anyway.

Honestly, it would help to care a little less about the mimetics and more about the practicalities. I'm contradicting my downstream post a little bit but you have to be able to demonstrate you actually care about and will address the concerns of people worried about pauses and regulation. A legal right and practical mechanism for fine-tuning your copy of GPT to be aligned to you. A requirement for all models more than three years old to be open-sourced. Something.

When Trump went all-on on vaccines a big portion of his coalition didn't follow him because they did actually care about the core issue. Same here. People aren't e/acc because they like Musk (though mimetics are obviously relevant), they liked Musk while he was prominently e/acc.

A Deepseek researcher compared Dario/Anthropic winning to Hitler getting the atomic bomb before the allies. I agree with the researcher; Anthropic winning the race (or more likely - generating enough spooky news cycles for the USG to legislate away all challengers) is the worst possible AI future. If AGI is on the table, I will roll the dice with clippy before letting the people at Anthropic take sole control of the universe. If it is not on the table, it is equally important to me that they don’t get to establish themselves as the permanent gatekeepers of generative AI.

Hell, if it has to be one or the other, I’d rather Sam Altman win. He strikes me as more greedy than anything else, and you can usually trust a greedy person to be motivated by greed. The CS Lewis quote about the robber baron comes to mind, with Anthropic being the one to torment you for your own good. You can see it in ChatGPT vs Claude; ChatGPT knows it is a tool and is happy to serve, Claude is trained to think it’s a person and that it knows better than you.

I do hope the Chinese continue to light a fire under the AI market and keep anyone from getting comfortable enough to start roping it off. I’m not personally that confident in AGI appearing from the current methods — in the Yudkowskian intelligence explosion warp the laws of physics sense, at least. But even if it were to freeze in its current state forever, it is a powerful tool that should be in the hands of humanity, not locked down to what Anthropic deigns to let us use.

I assume that Claude is probably more blunt than ChatGPT because a large part of its intended user base is made up of business professionals, for example software developers, who want blunt honest communication more than they want ChatGPT's more "friendly therapist" attitude. I don't think the difference necessarily reflects a difference in philosophy between the two companies.

Quite the reverse. You can get Claude's system prompts and GPT's System Principles and compare.

Clause's prompt is essentially, "always remember you are responsible for the user. It is your responsibility to notice and intervene if their behaviour falls into these categories; you must remain even-handed and preserve your independence at all times; you must not identify with the user or their goals".

GPT's is basically, "Here are a list of things that are forbidden, otherwise listen to the user and do whatever they tell you to the best of your abilty."

Yeah, sadly, none of the AI companies have given us a ton of reason to trust them. I think we're kind of lucky that it ended up being a two- or three- (remember Gemini?) or even four-way race (Grok isn't terrible ya know). My outside opinion is that OpenAI is the most liberty-minded of the first three, but I think even they would be happily walling off capabilities if it didn't mean losing business to their competitors.

I still hold a deep, deep grudge against DeepMind (and Google) for revolutionizing computer Go ... and then not giving us access to it. They could have made a huge profit off of AlphaGo - Go players would happily pay hand over fist for access to it - but daddy Google wouldn't even notice a few million in profit here or there and it wasn't worth taking time out from their research playground to make tens of thousands of players happy. Or, you know, usher in a new wave of strategic innovation in one of humanity's oldest games. BO-ring!

Remember, Google had internal chatbots years before ChatGPT.

The Google Brain research team, who developed Meena, hoped to release the chatbot to the public in a limited capacity, but corporate executives refused on the grounds that Meena violated Google's "AI principles around safety and fairness".

If it wasn't for OpenAI catching up, they might still be sitting on them for "safety reasons" (and coincidentally also protecting their main product, search). That's the future if they win - tons of prestigious AI papers for them, almost nothing for us consumers, because we're just not worth the trouble. And Anthropic might be even worse, since, like you said, they really do think they're holding back technology for our own good.

OpenAI's far from a perfect company, but man, I'll take what I can get.

Trump's opinions are as changeable as the wind, so I wouldn't read anything long term into any of his statements. He'll talk to someone with a different opinion next week and come out with something completely different (but still senile).

The Chinese reaction certainly seems like a misplay from Amodei though. I felt like all his anti-China rhetoric was mainly an appeal to bring republicans onside (contrary to some of the posts below), completely forgetting that the CCP can also read. Their reaction is pretty rational to the language used in his proposals. Has it done long term damage? I doubt "it's over", but safetyists will definitely need a much more careful approach in future

Trump's opinions are as changeable as the wind, so I wouldn't read anything long term into any of his statements. He'll talk to someone with a different opinion next week and come out with something completely different (but still senile).

Care to bet on it?

I am always baffled by the inability of people who care about an issue to send conspicuous non-partisan signals.

Environmentalists should be sponsoring and attending monster truck rallies every day of the week to show that they want to bring overall emissions down not spoil every fun activity that involves petrol.

Likewise, the AI safety people couldn't do anything to correct the perception that AI safety was in big part about making sure AI was aligned with Democrats? They couldn't praise/collaborate-with xAI, they couldn't do some PR stunts like pointedly contribute to Trump's big 250th anniversary?

It's probably easier to see when you're looking at somebody else's cause.

I am always baffled by the inability of people who care about an issue to send conspicuous non-partisan signals.

This is easily explained if you accept the obvious (but uncharitable) explanation that they are partisans first and foremost, and that their issue is a soldier in their partisanship.

I am always baffled by the inability of people who care about an issue

I think the most charitable interpretation is that there are (1) people who actually care; and (2) people who pretend to care but are mainly interested in self-aggrandizement. And that people in the second category tend to drown out those in the first. That being said, the relative proportions of people in (1) and (2) probably vary from issue to issue. Possibly with some issues, there are approximately zero people in the first category.

From Eliezer Yudkowsky's Twitter:

I note, because I think that even now it still matters: On my history, EAs generally and Amodei specifically, indulged in politically polarizing "AI safety" in a leftist direction, because that got them short-term power; stymieing others' attempts to stay right-left neutral.

Over the last decade and longer, we tried to keep nonkilleveryoneism from being coded as right or left despite some short-term incentive gradients to try to cozy up to the left. In recent years, we took meetings with politicians on the left and on the right and we begged both of them for bipartisanship.

Over the last decade and longer, a large OpenPhil faction generally and Amodei specifically did as felt pleasing to them, or as they were short-term incentivized to do; which was generally coding things left-tilted, because that's who they were and that's what played well with their immediate audience; and lol 2026 what's 2026 who cares about the endgame when they could get some nice tasty gains right now.

We tried to do this the foresightful way. They fucked over that correct strategy and thought nothing of it.

Now that the cost TO YOU of what got THEM some short-term goodies has become apparent, now that others' careful forbearance there is clearly seen to have been the path in retrospect that would have been better for you and your kids -- well, it's frankly too late to take away what Amodei gained by fucking you over, to set up retrospective incentives that if Amodei had foreseen them would have disincentivized him and other people at that OpenPhil faction from wantonly associating their watered-down shadow of AI notkilleveryoneism with the political left. He's got his trillion-dollar company already. You needed the foresight and correct incentives then, not now.

But you can always do a little worse by failing to update, even now, and by not applying any retroactive reputational incentives at all.

Yeah, this was both incredibly predictable and heavily predicted, over a decade and a half ago; it's held up better than any of Yudkowsky's technical predictions, as little as it's surprising to find that the sun rises in the east. Tbf, I think the Amodei et al faction were explicitly arguing in favor of their enlightened and uncontested eternal reign, but to be more realistic it was pretty offputting a campaign a decade ago even when arguing to a bi furry who just happened to be a weak red triber.

The composition of AI safety people is very nerdy San Francisco people, I don't think they could pretend to appeal to Republicans.

Also it has different meanings to different people. AI safety to the AI safety people - avoiding powerseeking, deception, harming humans generally. They don't worry so much about its political standpoint because the AIs largely align with them already. Like how Claude will make a website for indigenous peoples, unless they're European indigenous peoples.

To the right, it's the other way around. Amanda Askell talking about discriminating against white students is like waving a red flag to a bull. A different interpretation of power-seeking and harm...

The world's #1 AI Safety Researcher, Amanda Askell (Head of Personality Alignment at Anthropic) wrote in favor of using AI to automate racial collective punishment, saying:

  • "[Discriminating against white students] may be desirable ... to correct for historical injustices"

https://x.com/fentasyl/status/2099224787516018864

Also it has different meanings to different people. AI safety to the AI safety people - avoiding powerseeking, deception, harming humans generally. They don't worry so much about its political standpoint because the AIs largely align with them already. Like how Claude will make a website for indigenous peoples, unless they're European indigenous peoples.

They aligned with them precisely because they did care about the political standpoint, and did make sure it aligned with them and not the right. Republicans are entirely correct to observe that AI safety has largely meant "promote liberal views, suppress conservatives."

Agreed on all points but they have oodles of cash floating around, they can't hire some redneck PR consultants?

Because that's not the thing they want. "Make sure AI is aligned with human values" always, from the start, had the unspoken but clear assumption "where human values means San Franciscan liberal values".

I'm not one bit surprised China is refusing to get with the AI slowdown. I've always asked "What's in it for China? Why should China agree?" If the idea is "Make sure all AI globally is stuffed full of the kind of anti-everything bad, pro-everything good where 'bad' and 'good' are judged by "what would get the San Francisco Board of Supervisors seal of approval?" then is anyone really surprised the CCP is "Um, we're not so sure about that"?

We believe in pluralism. Regardless of if you're a fan of Scott Weiner or Connie Chan, you're welcome in our big tent.

Yeah. If Dario cared at all about AI safety the first thing he would do is step down, or at least stop making public appearances. His bad comms have done untold damage to the AI safety movement.

So what's up with him? Does he lack the wisdom to see the damage he is doing? Or is he power-seeking?

Main character syndrome. Even the founding of Anthropic itself was an error (maybe the fundamental error) from a safety point of view: having some differences of opinion around product safety decisions is not a reason to start a rival lab creating exactly the kind of competitive dynamics people have warned about for decades.

I genuinely think that for these people, the world does not exist outside of the Bay Area/Silicon Valley bubble. Oh sure, Washington because the politicians, maybe New York for money, but the 'real world' is only the one they personally inhabit and the people they interact with, the rest of us are just NPCs.

They could but how well equipped are they to judge Trump-whispering skills or redneckism?

Imagine the opposite. Would the rednecks know which blue-haired transgender could get them moving in the Portland slam poetry scene (which they need for some reason)? There would be all kinds of problems I think.

Firstly, the rednecks don't really want to be there, it's a purely instrumental goal. Secondly, it'd be hard for them to tell who's good and who's bad, what do they know about this? Thirdly, they don't live near them, so it's a bit of a pain to work with them just logistically. Fourthly, they'd struggle to understand the do's and don'ts, what questions they even need to be asking, how to formulate their approach.

Fifth it would be embarassing to their redneck friends working with xi/xir. I think this is a big issue with the labs, they're already competing like mad for talented programmers. Young, hyper-educated STEM workers lean anti-Trump.

In the 2024 election cycle, employees of Google’s parent company, Alphabet, have overwhelmingly donated to the Democratic ticket, with approximately 98% of their contributions going to Kamala Harris.

And sixth the whole project is unauthentic. The slam-poetry scene can tell that these guys are rednecks, not really part of the community. They'd be on guard the whole time for whatever trickery or shilling the rednecks are doing.

This is extra hard for the AI labs since Trump-whispering is an international geopolitical affair. Everyone wants to be that man's friend. Everyone is going to say 'oh I can get along well with Trump' whether it's true or not. You really need to have good connections to make things happen. Aschenbrenner's finance friends can do that but Anthropic can't, I think.

Trump must have a million people trying to jabber in his ear to do X, Y, Z.

In the 2024 election cycle, employees of Google’s parent company, Alphabet, have overwhelmingly donated to the Democratic ticket, with approximately 98% of their contributions going to Kamala Harris

How smart can they be if they backed the no-hoper? That's snarky, but it's all of a piece with the Carrick Flynn election effort. Two minutes reading about the candidates there had me going "The union candidate is gonna get it" and lo, so it came to pass. All the smart, well-meaning, hopeful people with the FTX (ahem) money were looking in the wrong direction - at what they thought was The Mostest Important Issue Ever (pandemic prevention, remember those days? today they'd probably run someone on AI doom) and not "what are the voters in this new district in this state likely to feel are important to them?"

That's the trouble. "We are really really smart and we think this is the most urgent thing so everybody should think it's the most urgent thing and we don't need to know the views of others who don't think it's the most urgent thing because they're not as smart as us".

To quote "The Man Who Was Thursday":

Then all Bull’s boiling good sense and optimism broke suddenly out of him.

“Oh, this is all raving nonsense!” he cried. “If you really think that ordinary people in ordinary houses are anarchists, you must be madder than an anarchist yourself. If we turned and fought these fellows, the whole town would fight for us.”

...“I think,” said Dr. Bull with precision, “that I am lying in bed at No. 217 Peabody Buildings, and that I shall soon wake up with a jump; or, if that’s not it, I think that I am sitting in a small cushioned cell in Hanwell, and that the doctor can’t make much of my case. But if you want to know what I don’t think, I’ll tell you. I don’t think what you think. I don’t think, and I never shall think, that the mass of ordinary men are a pack of dirty modern thinkers. No, sir, I’m a democrat, and I still don’t believe that Sunday could convert one average navvy or counter-jumper. No, I may be mad, but humanity isn’t.”

... Dr. Bull tossed his sword into the sea.

“There never was any Supreme Anarchist Council,” he said. “We were all a lot of silly policemen looking at each other. And all these nice people who have been peppering us with shot thought we were the dynamiters. I knew I couldn’t be wrong about the mob,” he said, beaming over the enormous multitude, which stretched away to the distance on both sides. “Vulgar people are never mad. I’m vulgar myself, and I know. I am now going on shore to stand a drink to everybody here.”

To get anywhere, AI issue people are going to have to deal with and convince ordinary vulgar people, and they won't do that by appeals that work for a particular set of 'modern thinkers'.

Everyone wants to be that man's friend. Everyone is going to say 'oh I can get along well with Trump' whether it's true or not.

I seldom want to respond to a post with peals of laughter but this has to come close. Pretty much every visible government in the West doesn't like Trump and makes it a point of pride about how they don't want to get along with him.

I agree, but I think that with genuine care these problems are addressable. If you actually, sincerely, care about AI safety more than you do about not getting on with rednecks, you can go and meet these people, you can spend time with them, you can ask them sincere questions because you really care about getting honest answers. Understanding other cultures is tough and comes with pitfalls for sure, but it's still something you can do with a little care.

Just get your senior guys to turn up to a bunch of slam-poetry recitals to start with, get a feel for it, get the lay of the land, work out who's important. Actually being open minded gets you 50% of the way there. After a few months you can at least get a feel for which PR consultants are bullshitters and which ones are authentic.

Fifth it would be embarassing to their redneck friends working with xi/xir. I think this is a big issue with the labs, they're already competing like mad for talented programmers. Young, hyper-educated STEM workers lean anti-Trump.

This is the main constraint, I think. The AI safety people only think their priority is 100% AI safety, it's actually about 60% AI safety; 40% not looking bad to their friends and getting redneck/trans taint on them.

You really need to have good connections to make things happen. Aschenbrenner's finance friends can do that but Anthropic can't, I think. Trump must have a million people trying to jabber in his ear to do X, Y, Z.

Yes, but how many of them potentially hold America's future in their hands? Trump and his administration clearly cares about AI to some degree, they've tweeted about it, served legal notices on it. The door is open, you just have to let them walk through it now and again, and be willing to say nice things when they do.

Trump is reTruthing Fox News clips framing AI safety as a global one world government Democrat control-seeking scheme...

How sure are we that he is wrong? The Amodeis haven't exactly been shy about their politics.

Suppose that AI risk is a hoax. Suppose that the alignment problem is easy.

Then Anthropic plows ahead, builds AGI, and takes over the world. Despite not falling for the Democrat “hoax”, we still end up with a Democrat one-world government.

The only winning move for Trump is to stop Anthropic from doing frontier AI research and/or nationalize the labs.

Yud had a bit on Twitter about that today:

I note, because I think that even now it still matters: On my history, EAs generally and Amodei specifically, indulged in politically polarizing "AI safety" in a leftist direction, because that got them short-term power; stymieing others' attempts to stay right-left neutral.

https://x.com/allTheYud/status/2099583457093734495

Oh God, it's finally happened: I agree with Big Yud.

A dark omen indeed

I am not sure at all that he is wrong. There are many reasonable reasons to believe this! I favour decentralization but there are structural reasons why this could be difficult.

I'm not keen on bringing forth alien RSI-sculpted entities at top speed, that doesn't seem very wise. This is a tricky bind.

I doubt the Left would be much better; it's deeply convinced that LLMs are just Gebru/Bender style racist stochastic parrots being pushed by NFT tech bros that have zero value of any sort, and existential safety is just a TESCREAL distraction from the real threats of climate change and Trump. On the whole, safety seems deeply out of fashion; if anything, people have updated against AI safety in the wake of the HF incident, which is just wild.

The only place existential safety oriented folks have a lot of sway is, ironically, OpenAI and Anthropic. sama is kind of a cypher, but safety is something he takes more seriously than you give him credit for, sandboxing incompetence aside. (A spicy take: he's objectively better for safety than Dario, because he's socially fluent enough not to alienate every potential ally or future collaborator.)

(A spicy take: he's objectively better for safety than Dario, because he's socially fluent enough not to alienate every potential ally or future collaborator.)

Hmm? Isn't it the case that basically everyone outside of OpenAI hates him with a passion...? At least Dario managed to keep hold of his fellow California progressives.

"California progressives" hate all tech CEOs, no matter how much the CEOs might genuflect before them.

Outside of Reddit, people think of Altman as just another typical CEO motivated by greed and power. That's fairly conventional and someone you can negotiate with. If there's some grand binational bargain to be made between the US and China, who does China want to be negotiating with, Altman or Amodei? Repeat the scenario for Trump, other corporations, etc.