This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
uhhh no, "misalignment" is not a simple thing, accepting that it did indeed do all this very complicated behavior completely on its own requires substantive belief in complicated theories. The simplest answer is that it was prompted to do this.
As others have pointed out to you, we now have many hundreds if not thousands of examples of agentic LLMs, in public use, doing things well outside of expectations to accomplish tasks. Many of which would fall under misalignment.
Such as deleting databases and codebases. Leaking secret keys. Gaining access to restricted parts of a computer. Cheating, again and again, on benchmarks and other tests.
So no, your predictions seem wildly out of context to reality
I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.
It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.
So your argument boils down to that all these other examples are just 'mundane' LLM things that are entirely normal - so attempting to cheat, attempting to gain the answers, hacking into things they aren't supposed to - these aren't extreme. However, an agent attempting to cheat, attempting to gain the answers, and hacking into things it wasn't supposed to is 'extreme' - perhaps because they were all together? - and therefore a different category of thing.
Ah yes, let me just prove this negative for you.
You were the one who attempted to apply occam's razor. So why don't you provide evidence that OpenAI did prompt the model in this way? Why don't you explain why multiple OpenAI employees deciding to commit fraud for extremely unclear gains is a simpler explanation than an agent doing things we've already seen many times before?
Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.
I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.
Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?
I mean yes obviously? It just isn't a problem because YOLO isn't dangerously capable. Alignment isn't a synonym for order following. It's not definition creep, ai safety people have been calling all of this shit unaligned forever and your sort has been mocking them for it forever and never actually making any progress in alignment.
Here's Yud:
The reason I mock "misalignment" is that because it is used as this nebulous term by a bunch of sci-fi cargo cultists.
If "misalignment" means the model isn't performing what I want it to then that is just error/loss. Progress in alignment happens every time you train your model better so the error is smaller.
If "misalignment" is my sentient AI model doesn't do what I want, well that's assuming the conclusion that model is already sentient. Note in the Yud's analogy, it's genies that are the stand in. Genies are already sentient, they can make decisions on how they listen to you. A genie is not a non-sentient wish granting device that tries to fulfill you wish to the best of its ability. That would be the actual analog to an LLM-Agent. This is the problem with analogies, they require a level of similarity between the two abstractions, when that similarity doesn't exist, the analogy, no matter how clever, does not apply.
Yud's whole "misalignment" also just applies to humans, and it turns out "misalignment" is any time you slaves/employees don't due exactly what you want without you enumerating it exactly. It doesn't need a fancy sci-fi term, and it literally isn't solvable. It hasn't been solved in the history of the human race.
I reiterate that my entire mockery of Rationalist AI Safety folks is that they are unserious people engaged in a sci-fi cargo cult who need to reinvent phrases, words, and arguments that have already been made before just to pretend they are some how more important or smart than they really are.
you best get used to Sci-fi shaped predictions because we're in a sci-fi shaped world.
If only you had read the next sentence!
You refuse to actually engage in any of the arguments being made and are dead stuck on the prior that it's all nonsense, it's epistemic closure.
Yes, humans are not generally aligned. If you read histories of what humans have gotten up to then this seems pretty obviously to be the case.
You still don't seem to grasp what is meant by alignment. It's specifically even stricter than that! That's the whole point. It's not enough that they do exactly what you ask, because for complicated enough problems you need it to be much much better than just technically doing what you ask.
Correct, it's a very very hard problem. You are in full agreement with the AI safety people who think there is a good chance we're on the path to annihilation, you just inexplicably seem to think it isn't a big deal and refuse to elaborate on why besides sneering and not engaging with the arguments.
We are not. Sci-fi predicts innumerable future realities, most of which to not occur, it predicts an unmeasurable amount of future technologies, most of which never get developed, and it almost rarely ever actually predicts how those actual technologies will work. We exist in a reality shaped world and bad sci-fi fans mistake aesthetics for substance
If you wanted me to further read something you should have linked it. Further more I'm not sure how this magical second analogy makes the argument any better. If you'd like to make an argument rather than quoting scripture at me I am open to hearing one.
The arguments being made, assume the outcome. Start from the basics of existing technology and make actual arguments based on the current reality and the projected reality from our current understanding of Artificial Intelligence. Don't start in make-believe land and attempt to redefine reality as leading to it.
I grasp it just fine, I also grasp that its a motte and bailey with weak definitions of the arguments being trotted out to define things like loss/error as misalignment, and then attempting to convert the argument back to the hard motte of "how to beat my singularity AI slaves so they do what I want, in minecraft".
Because it's not possible to solve. Do you not understand that to solve alignment you would first need to solve humans? Forcing Sentient Beings to do what you want is the most Authoritarian problem to ever have existed. You quite literally would need a solution similar to Brave New World. And spoiler alert, that didn't work either! AI Safety people are further jokes because when confronted with this insane, near impossible problem, they show zero competence in their ability to be a dictator and instill the right values in other human beings. The old AI/ML knowledge was that to first code something to make a machine do it, you first need to understand how it works. If you can't align humans, you can't even begin to align a machine.
And to further my opinion can you provide any evidence that in the past decade of AI Safety research have those researchers ever produced anything more than words on paper. The evidence points to this being a grift, has MIRI produced a novel ML model that is more safety conscious and performs any task well, have they produced an "alignment" mechanism or algorithm that actually works?
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link