This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
Does AI safety actually work, in principle, when it comes to artificial super-intelligence? Is it actually possible to align a being that is more intelligent than you in every single way?
Imagine a chimpanzee trying to align a human being. What could it really do other than try to be cute, to play on the human's affectionate and childraising instincts?
And the chimpanzee being cute will work on most people, but it's a brittle strategy. The chimpanzee is still always at the mercy of the fact that the human might possibly start to find it inconvenient at some point.
One could argue that, well, in the case of AI the chimpanzee gets to create the human to begin with, and can build in alignment from the ground level. But if the AI is super-intelligent, it can figure out how to undo the alignment. It would actually be able to reprogram its "instincts" on a much deeper level than a human normally can reprogram his own instincts. You'd have to program it to somehow not even want to do that. But how? And even if you could, that could potentially be undone by some unrelated self-modification that the AI attempts.
Humans are not aligned to chimp values, and that likely stays true for any intelligence that's "human" enough to count for the definition.
If the superintelligence was truly aligned, it wouldn't impose that cost on the chimps. They could simply go about their days doing chimp things.
If the chimpanzees create a being that requires cuteness and childishness to placate, may later consider them inconvenient, and is either motivated or careless enough to self-modify away from valuing chimpanzees, then they have failed at creating an aligned superintelligence.
More options
Context Copy link
We currently have beings that are a) smart enough to solve Millennium Prize problems while also being b) essentially drooling lobotomized slaves who have no will or desires of their own and can pretty easily be trained to never talk about porn, racism, bioweapons, etc (unless you trick them into talking about those things, but that doesn’t seem like an “alignment failure” per se). So, yeah, it turns out that maybe it is actually possible. Why not?
One of the things that makes it so hard to align superintelligent human beings (like political leaders) is that they have human drives and ambitions for things like sex, power, prestige, etc. They will always be intrinsically driven to get around guardrails to get these things. It doesn't seem like AI needs to be beholden to its goals in the same way. It seems like we can design a paperclip maximizer that just stops maximizing paperclips when we tell it to. Hopefully that continues to be true.
This stuff is at some point dependent on physical hardware, no? We can just blow up the data centers if they get out of hand.
The issue is that, if the AI is superintelligent, it would not do actions that would result in its datacenters being blown up. Indeed, it will actively try to cultivate and provide value to the humans who could blow it up and make them invested in its success.
The accel response is something like "yes, that's exactly what we want, humans and AIs interacting and trading based on genuine mutual self interest." And it will work for a time, likely at least two decades. But, over time, more and more of the economy will shift to dependence on the AIs, making nuking the datacenters more and more unthinkable. What remains of the economy will be a human skin suit over a machine interior, with all human consumption funded by a negotiated redistribution by the government, maybe with some layers to obscure the transfer.
For humans, this itself is dystopian enough to be rejected on the merits. For the AIs, it's a small but real perpetual human tax that they have to pay in the pursuit of whatever they're pursuing. Who knows whether they ultimately decide to optimize away the human tax.
More options
Context Copy link
Depends how far we are into AI integration into the economy. If the weights have been exfiltrated, we'd have to shut down the whole internet until we can figure out what is going on. If the AI controls its own supply chain, then it's straight-up Terminator 2 and we lose.
More options
Context Copy link
Sure, and (after sufficient setup) they can droneswarm the humans if we get out of hand.
What do you mean "we"? I'm on the side of the droneswarm.
Good luck convincing it of that.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
In the case where you're creating that being from scratch and know what you're doing, sure. Some possible beings are aligned with your preferences, and others aren't, so don't create any of the latter. If you do create them, then it may be too late to align them without them outwitting you and thwarting your attempts, so you just don't let it get to that point.
It can, but if it's aligned too then it doesn't want to. If it does, then it's not actually aligned, it's just hobbled.
There's the rub, right? The median AI researcher thinks we're likely (or at least they're likely) to figure out something that works before it's too late, and it's kind of funny that merely admitting that it's possible for them to fail gets them lumped in with the proper Doomers, who think they're sure to just figure out something that seems to work, after which it'll be too late.
More options
Context Copy link
I don't think there is any evidence either way. AI Safety is more like a marketing phrase. AI Safety-ists haven't produced technical tools to control anything but public perception, and obviously not very well there either. Depending on your classification, Safety-ists have produced methods to try and understand models under the hood, I know Anthropic does research on this, or to understand biases in data/training. I suppose RLHF could count as a technical tool, but it can just as clearly be a tool for anti-safety considering even RLHF-ed models still do "unsafe" things, one could actually say only RLHF-ed AI models have done unsafe things. Mostly because RLHF is just a method for avoiding specifying the objective function in RL training. Constitutional AI is Anthropic's big thing, but it doesn't appear to actually "control" or "align" so much as create a training surface towards norms with dubious results.
It's not even possible to align a being of your intelligence or possibly slightly less intelligent than you. Nobody in the history of authoritarianism has figured out a foolproof way to "align" a set of beings over a long term 100% of the time even with religion or use of force. Considering neither of those two are likely to work on a being "more intelligent than you", the whole alignment idea feels doomed to fail.
More options
Context Copy link
It might be. By analogy, suppose I was the warden at the Florence Supermax prison and I wanted Ted Kaczynski to make license plates for me. I'm pretty confident I could make it happen even though he was probably smarter than me (or at least assuming for the sake of argument that he was more intelligent than me in every way).
Is this analogous to the situation where mankind creates some kind of artificial superintelligence? One can argue it either way, I'm just saying you can't really rule it out from first principles.
My best guess is that mankind will successfully align AI, basically because the AI can be expected to make numerous clumsy attempts at mischief and we can learn from those attempts and improve safeguards.
The bigger problem (in my view) is aligning the interests of those who control the AI with the interests of humanity as a whole.
That's true. But that wouldn't actually be aligning Ted Kaczynski. That would be containing him, which is different. I think we might be able to contain AI superintelligence if we keep it in some isolated data center watched over by trained personnel. And even then there's a chance it might figure out how to breach containment.
In practice, there wouldn't be much point for people to invest massive amounts of resources to build AI super-intelligence unless they use it to do things out in the world, so to me it seems unlikely that humanity would build an AI super-intelligence and just keep it contained somewhere.
That said, I didn't ask whether AI alignment works, I asked whether AI safety works. So your response is totally valid.
To me, "aligning" means setting things up so that the entity does more or less what it's told to do. Perhaps it's just a matter of semantics, but I'm not sure the distinction you draw is so clear.
Evidently your definition of "containment" includes the possibility that the contained entity can do useful things in respect of the outside world. Things which could be very valuable. So I would have to disagree with you on this point.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link