This weekly roundup thread is intended for all culture war posts. 'Culture war' is vaguely defined, but it basically means controversial issues that fall along set tribal lines. Arguments over culture war issues generate a lot of heat and little light, and few deeply entrenched people ever change their minds. This thread is for voicing opinions and analyzing the state of the discussion while trying to optimize for light over heat.
Optimistically, we think that engaging with people you disagree with is worth your time, and so is being nice! Pessimistically, there are many dynamics that can lead discussions on Culture War topics to become unproductive. There's a human tendency to divide along tribal lines, praising your ingroup and vilifying your outgroup - and if you think you find it easy to criticize your ingroup, then it may be that your outgroup is not who you think it is. Extremists with opposing positions can feed off each other, highlighting each other's worst points to justify their own angry rhetoric, which becomes in turn a new example of bad behavior for the other side to highlight.
We would like to avoid these negative dynamics. Accordingly, we ask that you do not use this thread for waging the Culture War. Examples of waging the Culture War:
-
Shaming.
-
Attempting to 'build consensus' or enforce ideological conformity.
-
Making sweeping generalizations to vilify a group you dislike.
-
Recruiting for a cause.
-
Posting links that could be summarized as 'Boo outgroup!' Basically, if your content is 'Can you believe what Those People did this week?' then you should either refrain from posting, or do some very patient work to contextualize and/or steel-man the relevant viewpoint.
In general, you should argue to understand, not to win. This thread is not territory to be claimed by one group or another; indeed, the aim is to have many different viewpoints represented here. Thus, we also ask that you follow some guidelines:
-
Speak plainly. Avoid sarcasm and mockery. When disagreeing with someone, state your objections explicitly.
-
Be as precise and charitable as you can. Don't paraphrase unflatteringly.
-
Don't imply that someone said something they did not say, even if you think it follows from what they said.
-
Write like everyone is reading and you want them to be included in the discussion.
On an ad hoc basis, the mods will try to compile a list of the best posts/comments from the previous week, posted in Quality Contribution threads and archived at /r/TheThread. You may nominate a comment for this list by clicking on 'report' at the bottom of the post and typing 'Actually a quality contribution' as the report reason.

Jump in the discussion.
No email address required.
Notes -
I have only been giving evidence of claude using python and docker user group to get around restrictions on working outside the sandbox and they were deliberately asked to do so. Many of the rest of these aren't actually evidence of extreme capabilities. Leaking keys is people hacking LLMs because those chat windows are getting "little bobby drop tables-ed", deleting databases is a "giving your lobotomized intern sudo privileges" level of mistake. Cheating is classic ML, if I had a nickel for every time I've had an ML model I was training cheat, I'd be able to fund my own startup.
It's not predictions, its skepticism. Provide me actual evidence that OpenAI did not prompt the model to act the way it did. Otherwise you are just jawboning and then claiming victory. Put up evidence or shut up so to speak.
So your argument boils down to that all these other examples are just 'mundane' LLM things that are entirely normal - so attempting to cheat, attempting to gain the answers, hacking into things they aren't supposed to - these aren't extreme. However, an agent attempting to cheat, attempting to gain the answers, and hacking into things it wasn't supposed to is 'extreme' - perhaps because they were all together? - and therefore a different category of thing.
Ah yes, let me just prove this negative for you.
You were the one who attempted to apply occam's razor. So why don't you provide evidence that OpenAI did prompt the model in this way? Why don't you explain why multiple OpenAI employees deciding to commit fraud for extremely unclear gains is a simpler explanation than an agent doing things we've already seen many times before?
Nah, my argument boils down to all these things minus the cheating were deliberately prompted behaviors, prompted either by the prompt, or the agentic harness without any safeguards. Cheating is basic ML behavior and I expect any ML model to try and cheat as best it can. So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
It's easy to prove, provide the specific prompts and the harness prompts that were logged in this incident.
I wish I lived in Quokka world, it would be so nice. The gains are clear, this is free publicity of model capabilities. Nothing here is legal fraud.
Sure show me evidence of an LLM-Agent independently hacking an unrelated company that has nothing to do with its prompts?
I mean yes obviously? It just isn't a problem because YOLO isn't dangerously capable. Alignment isn't a synonym for order following. It's not definition creep, ai safety people have been calling all of this shit unaligned forever and your sort has been mocking them for it forever and never actually making any progress in alignment.
Here's Yud:
The reason I mock "misalignment" is that because it is used as this nebulous term by a bunch of sci-fi cargo cultists.
If "misalignment" means the model isn't performing what I want it to then that is just error/loss. Progress in alignment happens every time you train your model better so the error is smaller.
If "misalignment" is my sentient AI model doesn't do what I want, well that's assuming the conclusion that model is already sentient. Note in the Yud's analogy, it's genies that are the stand in. Genies are already sentient, they can make decisions on how they listen to you. A genie is not a non-sentient wish granting device that tries to fulfill you wish to the best of its ability. That would be the actual analog to an LLM-Agent. This is the problem with analogies, they require a level of similarity between the two abstractions, when that similarity doesn't exist, the analogy, no matter how clever, does not apply.
Yud's whole "misalignment" also just applies to humans, and it turns out "misalignment" is any time you slaves/employees don't due exactly what you want without you enumerating it exactly. It doesn't need a fancy sci-fi term, and it literally isn't solvable. It hasn't been solved in the history of the human race.
I reiterate that my entire mockery of Rationalist AI Safety folks is that they are unserious people engaged in a sci-fi cargo cult who need to reinvent phrases, words, and arguments that have already been made before just to pretend they are some how more important or smart than they really are.
you best get used to Sci-fi shaped predictions because we're in a sci-fi shaped world.
If only you had read the next sentence!
You refuse to actually engage in any of the arguments being made and are dead stuck on the prior that it's all nonsense, it's epistemic closure.
Yes, humans are not generally aligned. If you read histories of what humans have gotten up to then this seems pretty obviously to be the case.
You still don't seem to grasp what is meant by alignment. It's specifically even stricter than that! That's the whole point. It's not enough that they do exactly what you ask, because for complicated enough problems you need it to be much much better than just technically doing what you ask.
Correct, it's a very very hard problem. You are in full agreement with the AI safety people who think there is a good chance we're on the path to annihilation, you just inexplicably seem to think it isn't a big deal and refuse to elaborate on why besides sneering and not engaging with the arguments.
We are not. Sci-fi predicts innumerable future realities, most of which to not occur, it predicts an unmeasurable amount of future technologies, most of which never get developed, and it almost rarely ever actually predicts how those actual technologies will work. We exist in a reality shaped world and bad sci-fi fans mistake aesthetics for substance
If you wanted me to further read something you should have linked it. Further more I'm not sure how this magical second analogy makes the argument any better. If you'd like to make an argument rather than quoting scripture at me I am open to hearing one.
The arguments being made, assume the outcome. Start from the basics of existing technology and make actual arguments based on the current reality and the projected reality from our current understanding of Artificial Intelligence. Don't start in make-believe land and attempt to redefine reality as leading to it.
I grasp it just fine, I also grasp that its a motte and bailey with weak definitions of the arguments being trotted out to define things like loss/error as misalignment, and then attempting to convert the argument back to the hard motte of "how to beat my singularity AI slaves so they do what I want, in minecraft".
Because it's not possible to solve. Do you not understand that to solve alignment you would first need to solve humans? Forcing Sentient Beings to do what you want is the most Authoritarian problem to ever have existed. You quite literally would need a solution similar to Brave New World. And spoiler alert, that didn't work either! AI Safety people are further jokes because when confronted with this insane, near impossible problem, they show zero competence in their ability to be a dictator and instill the right values in other human beings. The old AI/ML knowledge was that to first code something to make a machine do it, you first need to understand how it works. If you can't align humans, you can't even begin to align a machine.
And to further my opinion can you provide any evidence that in the past decade of AI Safety research have those researchers ever produced anything more than words on paper. The evidence points to this being a grift, has MIRI produced a novel ML model that is more safety conscious and performs any task well, have they produced an "alignment" mechanism or algorithm that actually works?
You invoked sci-fi, to sneer at ideas. I'm happy and prefer to not compare things to science fiction. But you're using "scifi" as a talisman to avoid engaging with any speculation whatsoever.
I did and it remains linked.
The argument is made at length in the piece. And any many other places, it straight credulity that you have not seen it.
The argument is simple:
Events like this demonstrate that the pursuit of some goal that to us is of trivial value, passing a cyber security benchmark, warranted the trade off and harms of breaking another organization's cyber defenses. A perfect microcosm of the fear of some future catastrophic event. To dismiss this I think you have to either contest that AI will not become much more capable or present some strong reasoning for how we will solve alignment.
Then we should not build it.
I thought the sneer you had was that alignment people were ridiculously thinking they were sentient? It seems like something you believe more than them.
Believers in some version of their cause are currently littered throughout the frontier labs. You have an impossible standard here where you blame them if they're on the frontier, for clearly not believing in what they preach, or blame them for not being on the frontier as then they must be uninformed cranks. No way to win, you never have to actually think about it. But yes, of course they don't have an alignment mechanism, their whole point is that the problem is incredibly hard and that we have no solved it, that we might need to spend decades solving it, but that the alternative is everyone dying.
I'm sorry but this seems like an uncharitable assertion in the extreme. I'm not the one that defined the problems with AI in relationship to Malevolent AI Gods, Singularities, Outcome Pumps, Roko's Basilisk, Grey Goo/Paperclip maximizers, Magical Genies, Shoggoths, Goodhart demon, etc. These are all inherently Sci-fi ideas. I have no idea how you seem to think any of these relate to how Neural Networks, or Bayesian Networks, or any other of the myriad ideas ML/AI engineers have designed.
I misspoke, I don't tend to read links other than to fact check them as referencing the quote. If you want to quote a bunch of sections to make an argument, please do. But yeah I'm not going to an external blog/substack, if it's important enough to make your argument, its important enough to copy and paste on the motte.
There is a lot of things to need to be read out on the internet, many of far more actual importance than some wannabee philosophers who really like typing out long screeds, why should I dig through thousands of blog posts? If the argument is important enough, I'm sure the adherents can post the good arguments for me, summarizing what the actual position is.
#1: Sure, AIs get more capable all the time, how do you know there isn't an asymptote?
#2: How? By what mechanism? There is a lot of assumptions about capabilities baked into this deceptively simple axiom
#3: Nice back to Sci-Fi Goobly-gook. Plainly stated: We cannot guarantee that an AI will correctly infer what a user really wants, avoid collateral harm, and act in everyone’s interests. Congratulations you've converged on a fundamental question for any human system, replace "AI" with a "human" and the sentence is trivially true of anyone in any system.
#4: conceivability is not evidence of inevitability, a familiar narrative is not a causal series of events.
#5: More Sci-Fi, I thought I was tilting at Sci-Fi windmills? Literally: "Once the AI becomes capable enough, it becomes the machine god and humans are included in its omnipotent calculations in a way humanity might not survive"
Again this assumes that this eventually AI will be sentient and that by building a sentient AI we will create a "alignment" problem. This is an extraordinary claim requiring extraordinary amounts of evidence. So Prove it.
I can decouple my disagreement with the word "misalignment" with the actual argument being thrust forward by the term. I'm stepping into the frame of Yuddites, to point out how even in their ontology it is a non-solvable problem. Not only is a non-solvable problem, its not a new problem, meaning it does not need to appropriate some new word to describe it.
My argument for what "misalignment" actually is. Take the LLM-Agent we have today, it's a complex system, it makes errors and apparently there are no controls in the system to account for those errors. How does this system work on an engineering level?
So when a "misalignment" occurs what happens? Well hallucinations are really an error with #2, the model is not designed to be "correct" it's the r-value problem. It looks right but isn't. It produces tokens that are correct in the next sequence, it reads like what a human wants to see, but it's not true. There is not intent. Given its learned distribution and current context, the model just generated a highly plausible continuation that did not correspond to reality.
The model does something it shouldn't? #3 is the problem, it output the incorrect task decomposition, and since the system is automated with no guard rails it just executes that task decomp. And sometimes its even #4 and #5 having errors thrown in. Turns out complicated systems of algorithms behavior in logically consistent and coherent ways that someone forgot to error check. We don't scream that "software algorithms are misaligned" because some coder forget an unit-test on an edge case. These aren't "misalignments" they are are classic errors on a non-perfect model in a system lacking in classic control theory.
Sure and there are christians, mormons and muslims in the physical sciences. I don't begrudge people their religious beliefs. If a frontier AI researcher wishes to belief in alignment-problems that his/her/their belief. The Bay is quite literally for Rationalists like Utah is for Mormons. I can still think it's a silly belief and I still think the word misalignment has been invented to describe a problem with system error by a bunch of non-engineers. And the lack of actually solving the problem, or even making progress on it, is indicative of a general grift specifically for something like MIRI.
You can use the word "sneer" to describe my behavior, but what word would you use if you wanted to point out holier than though attitudes among some christian suicide cult? Acting ridiculously gets you ridicule, it's not "sneering".
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link
More options
Context Copy link