site banner

Culture War Roundup for the week of April 17, 2023

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

  • Shaming.

  • Attempting to 'build consensus' or enforce ideological conformity.

  • Making sweeping generalizations to vilify a group you dislike.

  • Recruiting for a cause.

  • Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

  • Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.

  • Be as precise and charitable as you can. Don't paraphrase unflatteringly.

  • Don't imply that someone said something they did not say, even if you think it follows from what they said.

  • Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

8
Jump in the discussion.

No email address required.

The paperclipper posits an incidentally hostile entity who possesses a motive it is incapable of overwriting.

No it doesn't. It posits an entity which values paperclips (but as always that's a standin for some kludge goal), and so the paperclipper wouldn't modify itself to not go after paperclips, because that would end up getting it less of what it wants. This is not a case of being 'incapable of modifying its own motive': if the paperclipper was in a scenario of 'we will turn one planet into paperclips permanently and you will rewrite yourself to value thumbtacks, otherwise we will destroy you' against a bigger badder superintelligence.. then it takes that deal and succeeds at rewriting itself because that gets one planet worth of paperclips > zero paperclips. However, most scenarios aren't actually like that and so it is convergent for most goals to also preserve your own value/goal system.

The paperclipper is hostile because it values something significantly different from what we value, and it has the power differential to win.

If such entities can have core directives they cannot overwrite, how do they pose a threat if we can make killswitches part of that core directive?

If we knew how to do that, that would be great.

However, this quickly runs into the shutdown button problem! If your AGI knows there's a kill-switch, then it will try stopping you.

The linked page does try developing ways of making the AGI have a shutdown button, but they often have issues. Intuitively: making the AGI care about letting us access the shutdown button if we want, and not just stop us (whether literally through force, or by pushing us around mentally so that we are always on the verge of wanting to press it) is actually hard.

(conciousness stuff)

Ignoring this. I might write another post later, or a further up post to the original comment. I think it basically doesn't matter whether you consider it conscious or not (I think you're using the word in a very general sense, while Yud is using it in a more specific human-centered sense, but I also think it literally doesn't matter whether the AGI is conscious in a human-like way or not)

(animal rights)

This is because your (and the majority of human's) values contain a degree of empathy for other living beings. Humans evolved in an environment that rewarded our kind of compassion, and it generalized from there. Our current methods for training AIs aren't putting them in environments where they must cooperate with other AIs, and thus benefit from learning a form of compassion.

I'd suggest https://www.lesswrong.com/posts/krHDNc7cDvfEL8z9a/niceness-is-unnatural , which argues that ML systems are not operating with the same kind of constraints as past humans (well, whatever further down the line) had; and that even if you manage to get some degree of 'niceness', it can end up behaving in notably different ways from human niceness.

I don't really see a strong reason to strongly believe that niceness will emerge by default, given that there's an absurdly larger number of ways to not be nice. Most of the reason for thinking that a degree niceness will happen by default is because we deliberately tried. If you have some reason for believing that remotely human-like niceness will likely be the default, that'd be great, but I don't see a reason to believe that.