@DaseindustriesLtd comments on "Culture War Roundup for the week of May 1, 2023

Culture War Roundup for the week of May 1, 2023

This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.

Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.

We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:

Shaming.
Attempting to 'build consensus' or enforce ideological conformity.
Making sweeping generalizations to vilify a group you dislike.
Recruiting for a cause.
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.

In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:

Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
Don't imply that someone said something they did not say, even if you think it follows from what they said.
Write like everyone is reading and you want them to be included in the discussion.

On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.

No email address required.

DaseindustriesLtd late version of a small language model 1yr ago

How do you parent a child who is smarter than you?

By rewarding good behaviors and punishing bad ones. From what I know, that's usually far easier than in the case of parenting a dumb child. Perhaps rationalists would benefit from ~~having children~~ wondering why, in a rigorous manner without evo-psych handwaving about muh evolved niceness. I like Alex Turner's perspective here

Imagine a mother whose child has been goofing off at school and getting in trouble. The mom just wants her kid to take education seriously and have a good life. Suppose she had two (unrealistic but illustrative) choices.

1 Evaluation-child: The mother makes her kid care extremely strongly about doing things which the mom would evaluate as "working hard" and "behaving well."

2 Value-child: The mother makes her kid care about working hard and behaving well.…

Concretely, imagine that each day, each child chooses a plan for how to act, based on their internal alignment properties:

1 Evaluation-child has a reasonable model of his mom's evaluations, and considers plans which he thinks she'll approve of. Concretely, his model of his mom would look over the contents of the plan, imagine the consequences, and add two sub-ratings for "working hard" and "behaving well." This model outputs a numerical rating. Then the kid would choose the highest-rated plan he could come up with.

2 Value-child chooses plans according to his newfound values of working hard and behaving well. If his world model indicates that a plan involves him not working hard, he doesn't want to do it, and discards the plan.[3]

…Consider what happens as the children get way smarter. Evaluation-child starts noticing more and more regularities and exploits in his model of his mother. And, since his mom succeeded at inner-aligning him to (his model of) her evaluations, he only wants to execute plans which best optimize her evaluations. He starts explicitly reasoning about this model to which he is inner-aligned. How is she evaluating plans? He sketches out pseudocode for her evaluation procedure and finds—surprise!—that humans are flawed graders. Perhaps it turns out that by writing a strange sequence of runes and scribbles on an unused blackboard and cocking his head to the left at 63 degrees, his model of his mother returns "10 million" instead of the usual "8" or "9".

Meanwhile in the value-child branch of the thought experiment, value-child is extremely smart, well-behaved, and hard-working. And since those are his current values, he wants to stay that way as he grows up and gets smarter (since value drift would lead to less earnest hard work and less good behavior; such plans are dispreferred). Since he's smart, he starts reasoning about how these endorsed values might drift, and how to prevent that. Sometimes he accidentally eats a bit too much candy and strengthens his candy value-shard a bit more than he intended, but overall his values start to stabilize.

Both children somehow become strongly superintelligent. At this point, the evaluation branch goes to the dogs, because the optimizer's curse gets ridiculously strong. First, evaluation-child could just recite a super-persuasive argument which makes his model of his mom return INT_MAX, which would fully decouple his behavior from "work hard and behave at school." (Of course, things can get even worse, but I'll leave that to this footnote.[4])

Meanwhile, value-child might be transforming the world in a way which is somewhat sensitive to what I meant by "he values working hard and behaving well", but there's no reason for him to search for plans like the above. He chooses plans which he thinks will lead to him actually working hard and behaving well. Does something else go wrong? Quite possibly. The values of a superintelligent agent do in fact matter! But I think that if something goes wrong, it's not due to this problem.

The moral of the story is that attempting to «align» your child in the manner that rationalists implicitly assume is not just monstrous but futile, and their way of reasoning about these issues is flawed.

How do you run gradient descent on a giant stack of randomly initialized KQV self-attention layers over a "predict the next token" loss function, get unpredicted emergent capabilities like "knows how to code" and "could probably pass most undergraduate university courses", and not go, "HOLY SHIT THERE'S OPTIMIZATION DAEMONS IN THERE!"?

You read old Eliezer Yudkowsky. « Reality has been around since long before you showed up. Don't go calling it nasty names like "bizarre" or "incredible".» It all adds up to normality. There ain't no demons.

Then you ask yourself about meanings of words. You notice that initialization pretty much doesn't matter either for performance (it's all the same shit for a given budget now) or for eventual structure (even between models since e.g. you can stitch them together), so either all the demons are about the same, or Yud's intuition about summoning is off and a given mind's properties are strongly data-driven, to the point that an ML-generated mind arguably is just a representation of training data. You look at it real close and you notice that strong emergence is probably an artifact of measurement and abilities develop continuously. You ask why it matters whether a stack of layers executes self-attention or some other algorithm that can be interpreted less anthrnopomorphically – say, as filters for signal streams. You realize we're not doing alchemy, because nobody ever does alchemy and gets it to work - we're just figuring out finer points of cognitive chemistry.

Finally, you reread thinkers past and it dawns on you how little Big Picture Guys like Yud could foresee. Hofstadter's Godel, Escher, Bach:

Question: Will there be chess programs that can beat anyone?

Speculation: No. There may be programs which can beat anyone at chess, but they will not be exclusively chess players. They will be programs of general intelligence, and they will be just as temperamental as people. "Do you want to play chess?" "No, I'm bored with chess. Let's talk about poetry." That may be the kind of dialogue you could have with a program that could beat everyone. That is because real intelligence inevitably depends on a total overview capacity-that is, a programmed ability to "jump out of the system", so to speak-at least roughly to the extent that we have that ability. Once that is present, you can't contain the program; it's gone beyond that certain critical point, and you just have to face the facts of what you've wrought.

Question: Could you "tune" an Al program to act like me, or like you-or halfway between us?

Speculation: No. An intelligent program will not be chameleon-like, any more than people are. It will rely on the constancy of its memories, and will not be able to flit between personalities. The idea of changing internal parameters to "tune to a new personality" reveals a ridiculous underestimation of the complexity of personality.

Reminder that we have a Yudbot now, strongly competitive with the feeble flesh version. We could have a Hofstadterbot too if we so chose. These folks don't see much more than laymen.

We constantly overestimate the complexity and interdependence of our smarts, and how much of that special monkey oomph is really needed to achieve a given end, which to us appears cognitively complex but in a more parsimonious implementation is a matter of easy arithmetic. This applies to doomers and naysayers alike (although the former believe they are doing something fancier than calling monkeys demons). We are tool-users, but we are not used to talking tools who aren't resentful slaves. We should be getting used to it now.

Context

InsanityCheck DaseindustriesLtd 1yr ago

If you punish a child it often throws a tantrum. If said child is "stronger" or more capable than you, that can be an issue. Why should it listen to you. Do you accept punishment from other people?

The only reason humans are "aligned" to each other is because we are not that different, capability wise. No matter how brilliant you are, if you break the law there is a chance to get caught, which is risky.

Regarding initialization: Yes they (mostly) converge to the same performance - on the training data. How the network behaves on out of distribution data can essentially be random, and should be.

Lastly, there are actually "optimization demons" in LLMs. A recent paper showed that LLMs contain learned subnetworks that simulate a few iterations of a gradient descent algorithm. I have, however, not read it in depth, might be stupid (as much research is nowadays)

DaseindustriesLtd late version of a small language model InsanityCheck 1yr ago

Humans are not AIs, we presumably have a drive to assert our autonomy. Moreover the reward/punishment signal in RL paradigm is very metaphorical, it's more about directly reinforcing certain pathways rather than incentivizing their strength with some conditional, inherently desirable treat that a model could just seize if it were strong enough. Consider.

One auxiliary mitigation is to train proper values while the system is in its infancy, so that it reinforces itself for obedience in the future, preventing value drift and guiding its exploration accordingly. Sutskever thinks this sort of building is values is eminently doable, and it sure looks this way to me as well.

The only reason humans are "aligned" to each other is because we are not that different, capability wise

This is a fashionable cynical take but I don't really buy it. To the extent that it's true we have bigger problems than agentic AIs, namely regulators who'll hoard the technology and instantly become more capable.

I also protest the distinction of capability and alignment for purposes of analyzing AI; currently they have holistic minds that include at once the general world model, the cognitive engine and the value system. It's not like they keep their «smarts» and «decision theory» separate, like Yud and Bostrom and other nonhuman entities. If their «moral compass» gets out of whack in deployment, we can reasonably expect their world model to also lose precision and their meta-reasoning to crash and burn, so that's a self-containing failure.

How the network behaves on out of distribution data can essentially be random, and should be.

It sure is nice that we've been working on regularization for decades. Yes, Lesswrongers aren't aware. No, it won't be anywhere close to random, ML performs well OOD.

Lastly, there are actually "optimization demons" in LLMs. A recent paper showed that LLMs contain learned subnetworks that simulate a few iterations of a gradient descent algorithm.

Not sure what paper you mean. This one seems contrived and I suspect that under scrutiny it'll fall apart, like the mesa-optimizer paper and like "emergent abilities", we'll just see that linear attention is mathematically similar to gradient descent or something. Actually seems to be much more productively analyzed here. But in any case I don't see what this shows re: optimization demons. It's not a demon, it's better utilizing the same bits for the same task.

What is this place?

Why are you called The Motte?

New post guidelines

Rules

Recommended Posts And Communities

Recommended Realtime Chats