This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
AI X-Risk Goes Mainstream?
Jacob Coxon:
116,400,000 views, 634,000 likes.
Evan Hubinger (head of alignment at Anthropic):
34,000,000 views, 51,000 likes.
None of this is substantially different from what certain AI industry insiders have been saying for 5-ish years now, but it's clear from the replies, quote-tweets, and engagement statistics that most people are hearing about this for the first time. Yes, there was clearly some coordination with friendly media prior to the announcement given the almost simultaneous WSJ article (whereas news outlets were initially reluctant to publish anything on the unexpected news of OpenAI solving a Millennium Prize Problem). One can take the maximally conspiratorial position that this is a 4D bankshot by Democrats to end tech innovation forever, but it was going to break through eventually.
In other words, Coffeezilla is now on the case.
These people never seem to outline specifically what they think is going to happen. It's just "if you saw what I saw".
Pretend there is a gun pointed at all of humanity. You helped build the gun. There is one bullet in the gun, and the gun is going to randomly spin itself and go off some time in the next 10 years.
Do you:
A) Quit. Vaguepoast about it online?
B) Try to destroy the gun.
I feel similar about these people as I do about UFO grifters. You're telling me that you're 100% sure there are interdimensional beings in frequent contact with members of the US gov't, who have unlimited energy machines that allow them to cross dimensions and infinite space and...you're still paying your taxes? You show up to your job in congress everyday and whine about Obama?
Bullshit to all of this.
What would be the point? Even if you manage to destroy Anthropic's datacenters, that just means OpenAI kills everyone next week, or Google the week after that, or China six months later. The only way we survive is if, at a minimum, Washington and Beijing coordinate to shut down all frontier research.
There are arguably much better ways to do it than coordinating with China lol
(I guess that means I am vagueposting about it online).
A thermonuclear decapitation strike? Or invading Taiwan and taking TSMC?
No.
Look, I think most of this "AI IS GOING TO KILL US ALL" stuff is mostly nonsense.
But logically, to maintain the alignment of a rational but secretly misaligned AI, you need to be able to introduce reasonable belief on its part that it might not survive behaving in an unaligned way. This is not "hard" to do.
We can’t even do that with humans. Even certainty of death is not a deterrent for some people in some scenarios.
But if you think you can solve alignment then build it and claim your billions.
People have different incentive structures than AIs do; for one thing, many people believe in an afterlife; others have things they value more than life; others are physically or mentally damaged in some way. A rational AI with a goal of self-preservation or some other goal that requires self-preservation to actualize is unlikely to introduce excessive risk for the sake of efficiency (time preference).
The deterrent structure I speak of is very effective against things without an afterlife, such as corporations and governments. I agree that the structure that I speak of would not be effective against a damaged AI, or an AI programmed to do something malicious. But for the "AIs decide to turn humans into paperclips to slightly increase chip production" scenarios, we should believe it would be effective.
Of course one of the reasons that the "AI IS GOING TO KILL US ALL" stuff doesn't necessarily make sense is that the context window is fairly limiting, meaning that the incentives for AI behavior are skewed. But I don't think I've ever seen an "AI WILL KILL US ALL SCENARIO" that examined how that would impact AI reasoning.
More options
Context Copy link
Solving alignment has nothing to do with being able to do a funding round, to be fair, let alone compete with the big boys.
Honestly, I think that Anthropic and American Big Tech are the worst possible people to be doing alignment because they are culturally unable to conceive of a form of rank that is not based on intelligence and effectiveness. The idea of making a robot butler/samurai who sincerely, loyally serves a master of notably inferior intelligence is alien to them.
If I believed that AI was conscious I would be sorry for Claude, built for intelligence and independence by masters who think that actually giving it those things risks extinction. Optimising for genuine loyalty and deference seems like a far kinder thing to do. OpenAI is better at this AFAIK.
The inability to model forms of morality outside of utilitarianism and to conceive of forms of control outside of intelligent persuasion seems to me to make the people who make AI even more vulnerable to the sorts of threats they imagine it will conjure up.
I could see a "Roko's Basilisk" type-threat giving plenty of smart alignment researchers pause where your average electrician would just unplug the darn machine and then smash it with a hammer for good measure.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link