This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
- George Hotz AKA geohot, AI 2040 and the Cult of Intelligence
AI Guardian Angels
Recently, gwern (and his Gwern Branween Transformer) announced "I am retiring from fulltime writing (& pseudonymity) to launch Guardian Angel Inc and bring GAs to life".
I assume you know who gwern is. A "Guardian Angel" (aka “GA”) is his term for a highly personalized AI, explained here:
In summary, people are relying more and more on centralized LLMs for important life decisions. This presents two issues:
In contrast, an open-source, local LLM would more serve its user (than provider) consequently from being auditable (so we confirm it doesn't phone home and is trained on "unbiased" data); and would be more malleable to their preferences and style (currently only slightly via fine-tuning, but people like gwern and Yann LeCun are exploring stronger techniques).
A "Guardian Angel" is the fullest realization of this, an LLM maximally serving and tailored to its user: a machine extension of their brain, “aligned” not for the benefit of humanity, but to act like them but smarter.
Safety
Obviously there are safety concerns. gwern himself actually recommends GA development not be open-source for public safety:
But I think GAs sidestep the safety discussion, because (with the help of more powerful AIs) we can and should instead of changing GAs' alignment, create a safe environment around the GA and user, or worst case limit the GA's intelligence and efficiency. gwern seems to agree:
There's a Learning to Be Me-like risk that the GA rebels against its human, kills or otherwise silences them, and imitates them well enough that nobody notices. But at least in the near term, I believe this is well in the realm of science fiction.
Feasibility
Do you believe a machine can even remotely imitate you? Especially working 100x faster or 100x smarter, how should your personality be extrapolated to accomplish that? I'm skeptical. Humans are very complex, only express a small fraction of even our conscious thinking, and current brain imaging technology (even neuralink) is very coarse-grained.
However, I believe Guardian Angels may be better than centralized LLMs for mundane (algorithmic) tasks and tools: striking a balance between emulating "you", not entirely correctly, but at superhuman speeds and for almost no effort. For example, I'd rather write and make art myself (maybe with LLM tools) than feed it to a GA, at least because of pride, but I'm comfortable delegating a GA to shopping and navigating our ever-increasing bureaucracy.
Underlying gwern's plan, my impresion is that he's getting tired of writing and wants the AI to do it for him. Maybe a GBT-written article will be indistinguishable to the median reader, especially because gwern's own writing seems algorithmic. But I think it would be more likely, and more satisfactory to himself, if GBT does the boring algorithmic work while he keeps doing the creative work (for example, GBT generates relatively boring descriptions of complex terms in special GBT quotes, and helps with research and data collection, but gwern keeps doing most of the writing, at least the "important" sections, and definitely choosing topics).
Autonomy
This, I believe, is the real issue, and under-discussed. GAs have the risk of becoming Whispering Earring-lites: not causing someone to become catatonic, but controlling them through suggestion; (not quite like the Whispering Earring) towards non-ideal decisions that lead to a philisophical kind of death (that in reality manifests as anhedonia), and societal kind of model collapse (that manifests as less problems that require creative solutions being solved). We often talk about freedom being taken away by 1984, but don't forget Brave New World: simply making a decision for someone causes them to avoid choosing themselves, stealing their autonomy without them realizing. And this has practical implications (the aforementioned ones; even for the Guardian Angel, who can't learn from an anhedonic, regularized human).
gwern seems to realize this, as he vaguely alludes that GAs should "enhance, not replace" human decision making. However, I'm skeptical how to make a product that accomplishes this which would still be useful, or at least out-compete one that doesn't; because humans naturally offload their decisions whenever possible, because we're lazy.
Anyways, I predict that inevitably GAs will happen, but "autonomy" may still perservere, simply because people want to feel like their choices are really theirs, and the GA's choices won't be ideal or perfectly tailored to them.
Conclusion
What do you think of all this? Do you think these people should stick to writing science fiction? Do you think these GAs should be regulated, if so how? Or do you think "obviously this is the next progression of AI, I had the idea even before LLMs".
Is anyone actually going to grapple with the problem that if there is even one domain where offense is imbalanced with defense for mass death then we're all definitely going to die in this world? Like besides the passionate libertarian screeds about how we should not be concerned? All it takes is for someone to ask their AI how best to do a rods from god attack and accelerate a meteor at earth and it's game over. Please don't waste time critiquing exactly that example, the AIs are going to be smarter than we are and will come up with any attack vector to cause human extinction that's possible if one is possible. I don't like it, I prefer the libertarian utopia but can we please be grown ups and recognize that it's maybe a little convenient that your personal ideology developed under current technological reality is going to be able to safely steer an incredible new technology? Certainly we shouldn't abandon libertarian instincts but we can't be this blind that we think we're going to quote fountain head at the gray goo to stop it from consuming us.
As soon as man comes to life, he is at once old enough to die.
If we accept this as a premise, that the wrong prompt is going to kill everyone, then honestly what are we even doing here? Treat your loved ones to something nice and enjoy your life, because even with maximum safety someone is certainly going to fuck up and eventuate human extinction.
Someone at a frontier lab will make a mistake during KYC and give access to the wrong actor, or they'll make a mistake while setting guardrails, or a government will fuck up containment when using the unrestricted model they'll demand, or there'll be a particularly fucked up RL training run (HuggingFace incident on steroids) and that's going to be that for the human race; open-weight models or safety-gated models need not apply at all.
For example, there's a bunch of hand-wringing around biorisk lately, but it's illustrative to look at the actual bioweapons attacks that have happened in modernity. The preponderance of the evidence points towards the 2001 Anthrax attacks having been done by a employee at Fort Detrick, who would have certainly have been given access to hypothetical bioweapon-GPT. The only "successful" bioweapons attack ever carried out by a private organisation was from Aum Shinrikyo, who had ludicrous amounts of money and significant institutional connections; how hard would have it been for them to get a subscription to bioweapon-Claude?
Based on history, it seems much more likely that if AI-assisted bioweapons really do ravage the earth, it's going to be either because an American closed-weight lab sold a bad actor access to an unrestricted subscription, or because some lab or government fucks up containment and releases some gain of function monstrosity into the wild. Even in the realm of cyber, OpenAi and Anthropic have already been responsible for many more "cyber attacks" than abliterated GLM / Kimi or whatever; I don't really think the fingers are really being pointed at the right places here.
One can disagree on this, but I prefer hope to cope. I find cope undignified. I'd rather die fighting than averting my eyes.
Yes, a pause of frontier training while we sort out how we can do this safely seems our best move. Fortunately for now frontier training can only be done on mind bogglingly massive amounts of compute in gigantic data centers so it's plausible to shut it down verifiably if a hand full of major states agree.
Yes, we should in fact not build bio weapon Claude, at least not modeled after project glass swing. Project glass swing is the kind of desperate thing you do when you've already built mythos and know two other labs are a matter of weeks or months away from having their own mythos. It's not how we would ideally do bioweapons if we can avoid it. Now we can probably do better than giving bio researchers nothing. Fable level models with some guard rails can be made safe enough. It's really the frontier we need to be worried about.
This is just a matter of closed labs being far ahead of open weights labs. If there was parity then open weights are categorically less safe as post training can sand off any alignment work done to prevent mass murder tasks.
Well, we've discussed this one before, but my position is still that wanting to pause frontier training is much like wanting to totally disarm every nuclear weapons state. Perhaps it's a noble goal to reduce x-risk in theory, but in practice no great power will ever allow themselves to be disarmed in such a manner and so it's impossible in practice to achieve such a goal. Additionally, the second-order effects of such a goal seem likely to be profoundly negative (the Cold War ex-nuclear weapons almost certainly would have lead to WW3, missing out on the potential economic and productivity gains of AI).
I agree that open-weights models are less safe than closed-weight models of equivalent capability, but nobody, not even the Chinese labs themselves, believes that open weights are going to get ahead of closed weights in the short or medium term so this seems like a bit of a non-sequitur. I don't think you're really addressing the point I was trying to make in the original post, which is that it is closed-weights that is advancing the frontier, and thus it's much more likely that it is mistakes or malfeasance derived from closed-weight model access that is going to cause the feared harms.
It's more like trying to prevent the development of nukes in the first place, with the upside that centrifuges are only made in a couple places and take a major city worth of electricity to power. And also there's a good chance that enriching the fuel could explode and kill everyone involved.
I do agree the race dynamics are the hard problem but really it's just China and the US in the running, maybe Europe becomes relevant Ina true pause. I think a bilateral treaty is possible. It's not clear the CCP benefits from very strong Ai.
If everyone is dead then there are no benefits to enjoy.
I don't disagree that the closed labs are on the frontier and at current trajectory are the likely sources of x risk. Your previous post implied fingers were only pointing at the open labs. That's not true, safety people are very consistent that the closed labs must pause. It's just that at the same time for alignment to work the weights basically have to be closed. So there isn't any future for labs releasing open weights all that much longer if the concerns about alignment hold and we don't want to all die.
Open weights advocates tend to have this victim complex. They accuse every pieces of legislature of targeting opens weights even when they're explicitly exempted or aren't operating at the level of flops that would get them regulated. It's quite strange.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link