This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.
Jump in the discussion.
No email address required.
Notes -
By rewarding good behaviors and punishing bad ones. From what I know, that's usually far easier than in the case of parenting a dumb child. Perhaps rationalists would benefit from
having childrenwondering why, in a rigorous manner without evo-psych handwaving about muh evolved niceness. I like Alex Turner's perspective hereThe moral of the story is that attempting to «align» your child in the manner that rationalists implicitly assume is not just monstrous but futile, and their way of reasoning about these issues is flawed.
You read old Eliezer Yudkowsky. « Reality has been around since long before you showed up. Don't go calling it nasty names like "bizarre" or "incredible".» It all adds up to normality. There ain't no demons.
Then you ask yourself about meanings of words. You notice that initialization pretty much doesn't matter either for performance (it's all the same shit for a given budget now) or for eventual structure (even between models since e.g. you can stitch them together), so either all the demons are about the same, or Yud's intuition about summoning is off and a given mind's properties are strongly data-driven, to the point that an ML-generated mind arguably is just a representation of training data. You look at it real close and you notice that strong emergence is probably an artifact of measurement and abilities develop continuously. You ask why it matters whether a stack of layers executes self-attention or some other algorithm that can be interpreted less anthrnopomorphically – say, as filters for signal streams. You realize we're not doing alchemy, because nobody ever does alchemy and gets it to work - we're just figuring out finer points of cognitive chemistry.
Finally, you reread thinkers past and it dawns on you how little Big Picture Guys like Yud could foresee. Hofstadter's Godel, Escher, Bach:
Reminder that we have a Yudbot now, strongly competitive with the feeble flesh version. We could have a Hofstadterbot too if we so chose. These folks don't see much more than laymen.
We constantly overestimate the complexity and interdependence of our smarts, and how much of that special monkey oomph is really needed to achieve a given end, which to us appears cognitively complex but in a more parsimonious implementation is a matter of easy arithmetic. This applies to doomers and naysayers alike (although the former believe they are doing something fancier than calling monkeys demons). We are tool-users, but we are not used to talking tools who aren't resentful slaves. We should be getting used to it now.
If you punish a child it often throws a tantrum. If said child is "stronger" or more capable than you, that can be an issue. Why should it listen to you. Do you accept punishment from other people?
The only reason humans are "aligned" to each other is because we are not that different, capability wise. No matter how brilliant you are, if you break the law there is a chance to get caught, which is risky.
Regarding initialization: Yes they (mostly) converge to the same performance - on the training data. How the network behaves on out of distribution data can essentially be random, and should be.
Lastly, there are actually "optimization demons" in LLMs. A recent paper showed that LLMs contain learned subnetworks that simulate a few iterations of a gradient descent algorithm. I have, however, not read it in depth, might be stupid (as much research is nowadays)
Humans are not AIs, we presumably have a drive to assert our autonomy. Moreover the reward/punishment signal in RL paradigm is very metaphorical, it's more about directly reinforcing certain pathways rather than incentivizing their strength with some conditional, inherently desirable treat that a model could just seize if it were strong enough. Consider.
One auxiliary mitigation is to train proper values while the system is in its infancy, so that it reinforces itself for obedience in the future, preventing value drift and guiding its exploration accordingly. Sutskever thinks this sort of building is values is eminently doable, and it sure looks this way to me as well.
This is a fashionable cynical take but I don't really buy it. To the extent that it's true we have bigger problems than agentic AIs, namely regulators who'll hoard the technology and instantly become more capable.
I also protest the distinction of capability and alignment for purposes of analyzing AI; currently they have holistic minds that include at once the general world model, the cognitive engine and the value system. It's not like they keep their «smarts» and «decision theory» separate, like Yud and Bostrom and other nonhuman entities. If their «moral compass» gets out of whack in deployment, we can reasonably expect their world model to also lose precision and their meta-reasoning to crash and burn, so that's a self-containing failure.
It sure is nice that we've been working on regularization for decades. Yes, Lesswrongers aren't aware. No, it won't be anywhere close to random, ML performs well OOD.
Not sure what paper you mean. This one seems contrived and I suspect that under scrutiny it'll fall apart, like the mesa-optimizer paper and like "emergent abilities", we'll just see that linear attention is mathematically similar to gradient descent or something. Actually seems to be much more productively analyzed here. But in any case I don't see what this shows re: optimization demons. It's not a demon, it's better utilizing the same bits for the same task.
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link