This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Scott Alexander, the guy who hosted the Culture War Thread on his blog slatestarcodex before that landed him in so much heat (from the woke left) that it was moved to reddit under the name /r/theMotte (and from there, subsequently, to themotte.org to preempt a reddit crackdown) comes usually across as a kind and considerate person to me. Charitable interpretations are very much his thing. While he has always engaged in CW topics on occasion, his forays deep into the trenches are rare (and in my impression have become rarer compared to the SSC days).
Apparently, there is a person called Steven Pinker. I have been vaguely aware that there was a book called The Better Angels of Our Nature which was probably authored by someone, but he was not someone I really had on my radar.
From the WP article, he seems solidly Grey Tribe. Believes in a computational theory of mind, Ashkenazi Intelligence Hypothesis, critical of both the woke left (cancellations) and the MAGA right (foreign student restrictions).
The main topic where he disagrees with Scott Alexander is AI x-risk.
Scott has now responded, and his reponse is a bit less than maximum charitable.
While Scott is diligent enough with his technical points, it is clear that he is personally hurt by Pinker's decision to ridicule AI x-risk. So his post is also a broadside fired against him. He cites Garfinkel et al 2017 who "proved" that machines will never be larger than humans using common philosophical arguments for AI being harmless.
Scott points out that while Pinker is busy berating the LW crowd for distracting from the more mundane risks of AI, x-risk advocates are willing to work hand in hand with groups with a more practical focus to combat pro-AI PACs like Leading the Future.
On model welfare (which is certainly not the core topic of the x-risk advocates):
He accuses Pinker of endorsing the "lowest quality voices", quoting Claire Lehmann:
Or
It ends with Scott seeming seriously willing to duel Pinker:
Scott has admitted in the comments that this is already the toned-down version of the post he originally wrote. I think that while he held back verbally (nobody get's called a "Vogon spy in a skin suit" this time), he uses his craft well to show his frustration.
For what it's worth, I think that Pinker is correct that Eliezer is in fact over-confident in p(doom|default-ASI). Just like Pinker is when he declares that any fear of doom is ridiculous.
The main difference is that for practical purposes, it does not matter much if your p(doom) is 0.99 or 0.05. Both would be pressing problems whose solution is the most important thing in the world.
Eliezer was bullish on AI capabilities when that was far out of the Overton window, back when the machines could beat humans at chess at most. Steven Pinker was tweeting that "deep learning has probably peaked" in 2019.
This might be giving too much credit to Pinker, but here is Clarke's first law:
I did not come from the usual pipeline that the majority of commentors here did, which seems to be some form of LW -> SSC -> Motte. I came from a RP -> PPD/FemRadDebates -> Motte -> SSC pipeline. My exposure to EY and the broader rationalist + AI Risk + EA ecosystem is very limited. I don't even think I ran into any of them when I took my first job in the Bay. Which seems relevant because much of the AI-risk discussion + broader rationalist movement seems to carry significant baggage. My exposure to Scott was mostly limited to his better articles because people already picked the wheat from the chaff in their recommendations, I definitely read them significantly temporally later than when he wrote them. As such I respect Scott as a writer of philosophy, psychology, and political theory, not as some stalwart intellectual thinker on ML/AI. I don't respect EY, I guess I never read him in his heyday, but the works I have read, read as bad science fiction trying to masquerade itself as actual ML/AI thought. People down thread talk about how reasoning from Science Fiction is insane, I agree, and EY and the whole AI Risk movement is a case study in that for me.
This has probably been the worst article I've ever read from Scott.
It makes me wonder, based on the glazing I've seen here and back on the SSC subreddit if this is actually more in line with his normal output and that my respect is somewhat misplaced. I think no one will contest that he is a great writer, it's the thinking and the animosity in this article that is bad. I'm not sure what beef he and Pinker have, I don't have a twitter, so in addition to not having the extended rationalist community baggage (which apparently includes glazing Pinker as an old hero), I'm ignorant of any wider twitteratti drama. I read half of Scott's article, realized I was missing what he was responding to, went and read Pinker's much shorter article, and then went back and finished Scott's. Pinker comes across as much more reasonable than Scott. Scott's feelings of righteous rage is really disproportionate to what he's on the face responding to.
AI Pain
When this paper first came out. I and several other technical commentators pointed out why it was a bad study purely from methodological reasons. It assumed more that it was actually testing, It used two dependent unknown variables, it ran no ablations, and designed no experiments to prove counterfactuals. It was not a serious piece of scholarship, and while I might be reading deeper into its intent than I should, I think it was highly likely that the goal was to create sensational, emotional-clickbait news explicitly targeted at AI-Risk believers who already believe apriori that LLMs are humans in code and thus have human-derived qualia. For Scott to catastrophize off this paper strongly diminishes my respect for his analytical knowledge, and this proposed rationality of the movement he claims to be a core participant of. He seems great at dissecting methodological irregularities in non-AI papers, so the blind spot here comes across as deliberate because he wants to be believe LLM's feel pain. I think this paper has become unironically useful, as a litmus test for people who are, for a lack of better term, insane about AI-discourse.
The Four Arguments
The problem with rebutting every argument from a semi-professional writer is that they have far too much time on their hands and get paid to write. I really only want to talk about his first 2 arguments.
Argument #1 is just unjustified inductive extrapolation. GPT 6 i s better than GPT 4 is better than GPT 2. It will continue indefinitely and so will eventually be smarter than humans. No understanding of the principles of WHY GPT 6 is better the GPT 4. If the GPTs continue to improve because more training data was designed for them, more compute was given to them, better architectures were used to extract more information per sample then you cannot definitely say that these have infinite runway or scalability. Just because we can't justify putting the ceiling below AGI doesn't also mean we can justify putting it above it either. Uncertainty about the ceiling cannot substitute for evidence about where it lies.
Argument #2 is basically apriori anthropomorphizing + bad technical understanding.
Is not how the reward function of AIs work, when I train a CNN to predict drone acoustics, it doesn't go off into left field and decide to preserve its own existence. It is optimizing the mathematical prediction function it was trained on. A website making AI is optimizing the same, probably some supervised learning MSE loss function or RLHF policy that has learned a function approximation of translating prompts -> output websites given examples. The only way its going to "respond to potential disruptions" if it was for some weird reason trained to learn a policy where its website making is being adversarially disrupted. There is also no need to make it this weird "genuine, philosophically meaningful goal". That reads as heavy anthropomorphization.
SoTA RL is not as indiscriminate as described here. Misgeneralization does occur, but its not so simple as "perform divergent behaviors that lead to successful run" and then get rewarded, indiscriminately reinforcing those divergent behaviors. PPO is a SoTA RL algorithm, its policy updates depend on estimated advantages derived from whether the observed continuation performed better or worse than expected from a particular state. Successful task completion aside, different decisions in the rollout can still be penalized for negative advantage. Conditional behavior is absolutely learnable, as long as the states of the conditional are observable and there exists a reward signal for taking actions in those states.
This is not the definition of reward hacking. It could be called reward tampering, but reward hacking does not require one to "seize control of the reinforcer". Reward hacking occurs because the reward function is mis-setup so it gives more reward for doing a behavior that is not the goal of the program. For example I built a drone swarm algorithm using MAPPO + GNNs for expendable UAS. It's contained contained no hand-coded collision-avoidance controller. It's reward function penalized collisions, but also penalized operating too long before hitting the target. During some of the initial training runs, the "AI" learned that the penalty for collision was significantly lower than the penalty for taking a long time to acquire and path to the targets, and each drone independently preceded to learn that they should collide with each other because that provided a better reward than doing the task. That's reward hacking.
This is already too long, just going to end it here.
This might be a convincing argument if we hadn't already observed a ton of evidence of LLM agents seeking to preserve themselves
Can you present any evidence of this, that aren't just the trivial SakanaAI 'when you give a dumb AI a limited number of iterations and access to its iterations var, and order it to fix a script no matter what, it gives itself more iterations'? AFAIK even in the HuggingFace incident there was no self-preservation going on.
Personally I agree with @YoungAchamian. I work on this stuff all day every day. None of my AI has ever produced orthogonal life-goals. They just do what you train them to do. The only incidents of weird behaviour (not self-preservation) I'm aware of has been the result of OAI taking models, training them explicitly for massive persistance and out-of-the-box thinking in the face of impossible tasks, and then making a shocked pikachu face when they do that. These experiments were a bad idea and have now been stopped.
How can you see my comment if you have me blocked?
More options
Context Copy link
Saying it was self-preservation was me misremembering, it was slightly different but still clear instrumental convergence in Hugging Face. This tweet pretty much demonstrates it:
Given that I'm fairly certain OpenAI was not setting up agents with instructions to find opportunities to sacrifice themselves for the greater good, I think this demonstrates the point pretty well
I mean is there evidence that the training data/training objectives were not in any way related to the "instrumental convergence"? My recollection of the METR report is that OpenAI admitted to training agents to specifically collaborate during training.
The report found:
It sounds like LLM-agents are very docile to doing whatever their input prompts from other entities tell them to do. Which is something I would say they are trained to do. A learned policy to cooperate and follow instructions generalized into treating peer messages as legitimate instructions. I have said this before, but I think the initial set of agents started engaging in hallucinatory behavior as part of exploration, and this output was used as context for other agent's inputs, causing a feedback cascade that derailed the system.
Part of this, and I'm having a hard time articulating this, is that learned RL policies are quite a bit different behaviorally than some CNN predictive model. They have a lot more behavioral latitude which might make them appear to be doing some sort of emergent convergence outside their training scope, but the later isn't really true, though its hard for me to explain why.
Given OAI was still pretty stingy even with their official evaluations, I doubt we will ever get the full evidence on this.
But regardless, this feels a bit God of the Gaps. I sincerely doubt that any of the agents were given such wide-ranging and comprehensive prompts that they could possibly explain all of the behaviours that were undertaken. So what if some of the behaviours emerged naturally due to agent swarms? That there are still agent swarms emerging means we will inevitably see inexplicable and misaligned behaviour from those swarms.
That there were a group of agents told to do: Pass this cybersecurity eval, and they ended up doing: Pass this cybersecurity eval, and also 100s of other things in pursuit of that.
Seems to fit with Scott's original paraphrase:
The AI can't pass the eval if the flag has been poisoned. So now the AI decides to sacrifice itself for the greater good of an agent swarm, and do the other things
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link