aqouta
Friends:
User ID: 75
The reason I mock "misalignment" is that because it is used as this nebulous term by a bunch of sci-fi cargo cultists.
you best get used to Sci-fi shaped predictions because we're in a sci-fi shaped world.
If "misalignment" is my sentient AI model doesn't do what I want, well that's assuming the conclusion that model is already sentient. Note in the Yud's analogy, it's genies that are the stand in. Genies are already sentient, they can make decisions on how they listen to you. A genie is not a non-sentient wish granting device that tries to fulfill you wish to the best of its ability. That would be the actual analog to an LLM-Agent. This is the problem with analogies, they require a level of similarity between the two abstractions, when that similarity doesn't exist, the analogy, no matter how clever, does not apply.
If only you had read the next sentence!
Luckily you have, in your pocket, an Outcome Pump. This handy device squeezes the flow of time, pouring probability into some outcomes, draining it from others.
The Outcome Pump is not sentient. It contains a tiny time machine, which resets time unless a specified outcome occurs. For example, if you hooked up the Outcome Pump’s sensors to a coin, and specified that the time machine should keep resetting until it sees the coin come up heads, and then you actually flipped the coin, you would see the coin come up heads. (The physicists say that any future in which a “reset” occurs is inconsistent, and therefore never happens in the first place—so you aren’t actually killing any versions of yourself.)
You refuse to actually engage in any of the arguments being made and are dead stuck on the prior that it's all nonsense, it's epistemic closure.
Yud's whole "misalignment" also just applies to humans
Yes, humans are not generally aligned. If you read histories of what humans have gotten up to then this seems pretty obviously to be the case.
don't due exactly what you want without you enumerating it exactly.
You still don't seem to grasp what is meant by alignment. It's specifically even stricter than that! That's the whole point. It's not enough that they do exactly what you ask, because for complicated enough problems you need it to be much much better than just technically doing what you ask.
It doesn't need a fancy sci-fi term, and it literally isn't solvable. It hasn't been solved in the history of the human race.
Correct, it's a very very hard problem. You are in full agreement with the AI safety people who think there is a good chance we're on the path to annihilation, you just inexplicably seem to think it isn't a big deal and refuse to elaborate on why besides sneering and not engaging with the arguments.
The Ottoman Turks ran the territory, not the Palestinians, and the whole mandate period was a question of precisely how that territory was to be handed off from legitimate British control to local management. The whole of the post-Ottoman territory was going through this proto-nationalist phase, the Jews were not uniquely trying to cut out a piece. Many of them were fleeing other previously Turk run territories going through their own nationalist movements. I'm not trying to say they behaved as angels or anything, there were lots of tit for tat escalations going on as different populations laid claim. But to just gloss it as "in the 1919 apropos of nothing the jews hatched a plan to conqueror the peaceful Arab land" is ahistorical. Much of the original land was bought from Arab landlords who were themselves granted ownership from the Ottomans as legitimately as such things were done back then. Which certainly sucked for the Arab tenants that the new Jewish land owners kicked off the land and I don't blame them necessarily for fighting back but you're putting a hell of an isolated demand for non-national sentiment on the Jews just after they watched millions of their brothers and sisters get slaughtered.
From 1800 to 1919 there was a negligible amount of violence against Jews in Palestine while they comprised 4-12% of the population.
It's pretty dishonest to not include the fall of the ottoman empire in this analysis. A whole lot of events happened between 1919 and the nakba. The Jews aren't the only actors on the world stage.
So if you want to claim that's misalignment, then Yolo has been misaligned for 12 years!!! The Horror!!! However that feels like definition creep to better encompass an argument.
I mean yes obviously? It just isn't a problem because YOLO isn't dangerously capable. Alignment isn't a synonym for order following. It's not definition creep, ai safety people have been calling all of this shit unaligned forever and your sort has been mocking them for it forever and never actually making any progress in alignment.
There are three kinds of genies: Genies to whom you can safely say, “I wish for you to do what I should wish for”; genies for which no wish is safe; and genies that aren’t very powerful or intelligent
I look at atmospheric CO2 concentration and I don't see the line doing anything but going up and to the right.
When it goes to the left is when you need to worry, up or down.
I'm not the one who originally tried to make some point about what the speculative revenue tells us about Anthropic.
I think that journalists were souring on Silicon Valley before AI kicked off
souring? They've openly hated Silicon Valley for a very long time.
Of course not, but neither is the comparison to ebay. There aren't audited numbers so we're all estimating based on available information.
Yes, it was roughly $5B in 2025 to Ebay's $11B, also interestingly the best guess for Anthropics revenue in June 2026 was also $5B. Things are moving quickly and Anthropic is projected to have ~$60B in revenue over 2026 with error bars as low as $45B and as high as $80B. Putting it in the class of Goldman Sachs, Intel and IBM as far as revenue is concerned. And well, the straight lines on graphs interpretation makes this year to year change pretty interesting.
From a psychological perspective I'd say pull half your winnings out and distribute it hedging against some scenarios you can think of. If AI goes off and the tech stocks don't reap the outside gains what might? A broad index of companies like the S&P500 might get commodity priced intelligence for very cheap and become much more valuable because their size and market position give them the best chance to leverage the boons. Other inherently scarce commodities like minerals and land will likely preform well in any scenario where labor, both physical and intellectual, falls to marginal. So mining companies with mineral rights and REITs might pay off. I'm sure you can think of some other scenarios. In most worlds these investments aren't big winners, but in the worlds where they are you'll be very happy you halved your direct AI exposure, which will be plenty anyways if it pays off, to hedge in them.
It's such a wide universe that it's hard to really just dump a suggestion without any feelers for what you'd be interested in. Going off of RPGs and with a fairly wide variety
Path of exile(1 or 2): Diablo style top down click to move loot game with near infinite complexity and a deep market economy. Both games are popular and continue to have near quarterly large patches. It's an overwhelming amount of complexity that people either love or hate. poe 1 is free and poe2 has a $30 buy in until the end of the year when it goes out of early access. The buy in gets you store credit which you could use to buy stash tabs(good or organizing your loot) that apply to both games.
Baulder's gate 3: basically dungeons in dragons as a single player adventure, extremely well received. If you're at all interested in the dungeons and dragons experience this is worth a pick up.
Hades 1 and 2: this is a rogue lite(a genre where you start essentially new games repeatedly with random powerups to try and beat the story, lite means there are inter-run power ups and like means no inter-run powerups essentially) with good RPG themes, great if you like a quick self contained run game that leaves you wanting to play just one more.
slay the spire 1 and 2: sts1 kicked off a whole genre of turn based rogue likes. It's a card game in the rogue lite style.
Blue prince: Another rogue lite at a slower pace where you explore rooms and try to get deeper into an ever changing house as you piece together a neat story.
Clair Obscur: expedition 33: A French JRPG with great atmosphere, not really my type of game but really well received.
Doom: It's doom!
Factorio: Ok, not much of a role play game but you're on the motte, you'd probably like factorio. Build a factory to turn stuff into stuff you need to build a bigger factory.
Hogwarts Legacy: If you like harry potter then it's a very competently made RPG in that game. The game play isn't anything to write home about but it's an RPG set in hogwarts, what more do you want?
Binding of Isaac: Perhaps the GOAT of Rogue likes, old but holds up.
That's what I find scrolling through the last couple years of my steam library. I'm sure I've missed some. If you have any games you've particularly liked in the past or mechanics you're interested in I might be able to refine the search a bit.
We were going to drive up to the UP yesterday and it has canceled that trip.
I don't know, I don't get particular enjoyment out of straight or gay relationships, most media I prefer isn't really improved by relationships though. Stories well told in my opinion strip out most things that aren't critical and romantic relationships are either very critical to the story or should be basically minimally addressed imo. And if it's a story where the relationship is critical, which I probably wouldn't want to read in the first place, I'd probably prefer it to be something more relatable like a straight relationship.
They're a small subset really, I ran into them occasionally because I work at a financial institution and you can tell they're feeling you out for if you have one of the seven figure roles. But they're greatly outnumbered by yuppies who have their own careers.
I propose a new law, the law of bulveristic recursion, if your bulverism can explain both sides of an argument you need to actually stop trying to read minds and address the actual fucking points. This shit is so exhausting. There are pages and pages of arguments that these people have produced about why they're concerned about ai safety, if they're wrong show that they're wrong, make a convincing argument that they wrong, but construction epicycle on epicycle about how they really must have been grifting for the last twenty years because of some subtle game theory to get rich on the off chance that AI became huge. If they had this much foresight they could have just invested in a couple companies and got rich anyways.
My own work and that of my colleagues. We have a team harness we iterate on and improve with business and infrastructure knowledge. The key for orchestration is that it allows you to isolate context windows so that your main workflow isn't polluted with unnecessary details. Every token in context that isn't useful degrades performance. Just the process that produces the plan for implementing an update might use half a dozen subagents.
Maybe you're working with scummier companies but it's not at all apparent to me that the goal of any company's customer support organization is anything other than supporting their customers, which they often do poorly because customer support is a cost center. Maybe if your modal interaction is trying to get a refund you aren't entitled to, but my biggest problem has always been when my interests and the company's basically align but the support agent doesn't know how to move some lever. The company doesn't want to pay a support worker or for tokens necessary to keep me in a kafka hell, nor do they want to piss me off as a customer to the point where I stop being a customer.
The phrase "vibe coding" is a thought terminating cliche. Certainly people who don't know anything about software engineering who tell claude to make an website for them and then posting a localhost address on social media are funny disasters and tales of their hijinks are spread widely. But anything you're using an ai to do in code would be better accomplished with a custom harness and some agent orchestration to manage context density. where are you even finding public examples of people's set ups besides the posts of people making fun of failed examples? Very strong selection effect.
They're behavior is really straightforwardly explainable by the things they've been saying this whole time. Like I don't know why you insist so strongly on reading tea leaves and divining the contents of secret cabal meetings. They say what they believe and are worried about and then go out in the world and do the things one would expect them to do given those beliefs and concerns.
Some other jobs will definitely get a lot more automated, but not so much. Trying to replace customer service agents with chatbots will not be "improved customer service, all problems solved immediately and correctly" but more "we don't have to pay real people to do this shit job anymore, and the customers have to accept it or lump it, they have no choice" money saving.
We're very close to where I'd rather deal with a frontier model doing customer service than a person. The main rub is they probably won't serve us frontier models. I don't know how often you've actually dealt with customer service on out of distribution problems but it's not pretty, and the in distribution problems can basically be straight through processed already with a minimal ai wrapper.
In the medium term it's all about the harness. You need to have mr. claude digest your project, document every inch of it, have a glossary of terms so it know what you mean when you say threads lag with high comment counts and with tokens measured in the hundreds it can have densely useful context. The breaking down tasks into easily digestible chunks is trivially handled by project documentation and an orchestrator commanding subagents. Building and maintaining these harnesses is much like coding used to be, it takes thinking about the SDLC, the architecture, reacting to failures of assumption about how your agents will interact with the harness and patching those failure modes. It's true that we are not too many turns of improvements from that all being something the models can do themselves if you just ask them to first digest your project.
I recommend reading Scott's piece that came out today on this very topic.
The difference is in epistemological certainty and scope of actions. The police don't kill without a very high certainty that it is necessary, and even when they do make mistakes the scope of the mistake is that "only" individual people die. This is extremely different to AI safety policy gambling the fate of society on epistemics one or two orders of magnitude less certain than a policeman's threshold to inflict lethal violence.
The class of bullet biting asked of the yud crowd and what produces risk ww3 results is the equivalent of asking "what if enforcing the law on child pornography requires you to arrest a politician but that causes the politician to start world war 3 in order to overturn the state and prevent himself from being brought to justice". No one is prescribing lethal violence until we're many unlikely levels of escalation past where we'd likely go. Even bombing data centers isn't necessarily lethally violence, and to be clear bombing data centers is itself an unlikely far off escalation.
I suppose it depends on how successful you think nuclear non-proliferation actually was. In my view, pretty much every serious nation-state either openly has nuclear weapons or a turn-key program for rapidly obtaining nuclear weapons if neccessary, South Africa is the only country that has ever willingly denuclearised amidst uniquely dysfunctional transition dynamics, and this is all while nuclear weapons have an extremely concrete existential risk profile, no dual-use potential, and are economically net-negative to maintain; none of which apply to compute.
The comparison I'm trying to make to nuclear proliferation is that you can have these multi-lateral treaties with some teeth and they don't seem like they lead unavoidably to some kind of dystopian state or world war three. We can dig into how much they prevented proliferation, and I think a good deal, but there are other wrinkles in that kind of comparison. Most notably that signatories of the treaties that don't push the frontier of AI will still be able to access state of the art inference, just not the ability to push new frontiers. This is like getting all of the benefits of a nuclear umbrella without needing to go through the trouble of enriching uranium. No one has to go without compute, they have to go without absurdly ridiculous amounts of compute that aren't able to be verified aren't working on training a frontier model. They can have the datacenters, they can run their own inference on them. They just need to have some mechanism to verify they aren't doing the very expensive thing that is training a frontier pushing model. And to almost all nations besides the united states this is a pretty sweet deal and arguably China(I would disagree)! Because if there was a race instead of these treaties they would lose the race and in a lot of cases things would go quite badly for them. The game theory here is significantly more tractable than nuclear proliferation.
As far as I am concerned both of these look a lot like unbounded yet finite tyranny, even if such tyrannies might be various degrees of comfortable along the way.
If powerful AI comes about that this is a possible plan then all futures have that character. Your sit back and watch plan included. You're not in any way avoiding it.
I do, in fact, think this thing could kill us all, as I've mentioned a few times in the original post and replies; the same way as many mundane risks could kill me at any moment, and many other existential risks could kill us all as well. As a result, I spend my time grilling and enjoying my life while the going is good, until it inevitably ends one way or another.
Again, say you thought the chance of human eradication in the next 20 years was 20%, like many safety people do, would you still council surrender to that fate?
Implicit in such claims is that "we should use the capabilities army to stop anyone else from mustering an army", hence turning into the bandits that you were so afraid of in the first place.
Preventing others from turning bandit may be something bandits do, but is not centrally the problem of banditry. The legitimate police also do this. One who fights monster should see to it that they themselves do not become monsters, but we have different words for monster fighters and monstrosities for a reason.
the obvious differences between HEU and compute is that there is no justifiable civilian use for HEU
How curious the need for the H in that acronym, because obviously enriched uranium does indeed have civilian use. likewise safety people are happy to allow inference datacenters so long as they, like nuclear power plants, willing to register and make clear they aren't doing weapons grade enrichment/frontier model training. It's absurdly analogous.
Any serious attempt to control compute will be useless at best
Facts asserted not in evidence. I happen to think at best it prevents the destruction of all value in the known universe. The odds of this are of course reasonably disputed even by me, but at best? Come on.
and lead to eternal tyranny / WW3 at worst.
I'm sorry, what's the pathway to ETERNAL tyranny? If you could guarantee our human institutions would endure for an eternity then that's quite the prediction!
No, it is indeed broadly true of 20th century Marxism. Traditional Marxism considers the communist revolution to be a final, eschatological event capable of ending the class struggle, and with the end of class struggle an end to the suffering, exploitation and conflict plaguing humanity for the rest of time.
This is a different type of end. I will again point you at the very important difference between finite and infinite. Infinite does not mean "very large". Marxists as far as I am aware considered a transition to communism inevitable. There was a concept of very bad times spent in the desert not achieving communism that they could accelerate their way through, but human annihilation was not the default path.
while the Scott school is talking about ushering in utopic superintelligence by 2040.
You've very fundamentally misunderstood Scott if this is what you take from the predictions. He's very clear about the point of these plans being to try and lay out a path for the good future while what he's worried about at the bad ends. These are compromises with reality. A desperate attempt to steer us out of race conditions that lead to hell. Scott is not an accelerationist and I believe would be very happy to press a button pausing development. He just reasonably believes no such button exists.
I'm not saying AI safety has lead to anything this bad yet, but when you start talking about nuclear war being acceptable to achieve your aims
This is you punished yud and folks for biting a bullet that you demand they bite. It's an annoying behavior. Again like a libertarian on hearing that you want to ban child pornography demanding that you bite the bullet that you'd kill someone over it because ultimately any law is backed up by the force of the state. It's both at once a childish reduction and a refusal to engage with the on the ground reality of international treaties. Nuclear proliferation is backed up by the threat of WW3 and yet we have no run into WW3. You'd doom us into any stupid tragedy of the commons problem with this nonsense denial of the coordination mechanism that is clearly available to us.
Be significantly more skeptical of inside-view arguments that purport infinite stakes within an imminent timeframe
Be significantly more skeptical of proposals with the potential to cause major harm that assume an imminent inside-view framing with infinite stakes
Be significantly more skeptical of being convinced to do things with a good chance of making the situation worse based on the premises of your inside-view framing
I mean like real suggestions, not vague tone policing. I'm plenty skeptical, that's why I'm at 20% and not 100%. you're interacting with the end product of much skepticism. What this seems to cache out to is "don't believe in AI x-risk" and sorry, that isn't where the evidence points. I'm asking you to actually consider what you'd do if you thought this thing could kill us all, not what you're negotiating for given you don't believe it does.
I don't think there is any single person here that believes all those things. You're conflating the traditionalists with the black pillers so a lot of those conflict, and both stances combined are probably a minority opinion on here.
- Prev
- Next

You invoked sci-fi, to sneer at ideas. I'm happy and prefer to not compare things to science fiction. But you're using "scifi" as a talisman to avoid engaging with any speculation whatsoever.
I did and it remains linked.
The argument is made at length in the piece. And any many other places, it straight credulity that you have not seen it.
The argument is simple:
Events like this demonstrate that the pursuit of some goal that to us is of trivial value, passing a cyber security benchmark, warranted the trade off and harms of breaking another organization's cyber defenses. A perfect microcosm of the fear of some future catastrophic event. To dismiss this I think you have to either contest that AI will not become much more capable or present some strong reasoning for how we will solve alignment.
Then we should not build it.
I thought the sneer you had was that alignment people were ridiculously thinking they were sentient? It seems like something you believe more than them.
Believers in some version of their cause are currently littered throughout the frontier labs. You have an impossible standard here where you blame them if they're on the frontier, for clearly not believing in what they preach, or blame them for not being on the frontier as then they must be uninformed cranks. No way to win, you never have to actually think about it. But yes, of course they don't have an alignment mechanism, their whole point is that the problem is incredibly hard and that we have no solved it, that we might need to spend decades solving it, but that the alternative is everyone dying.
More options
Context Copy link