@MathWizard's banner p

MathWizard

Good things are good

0 followers   follows 0 users  
joined 2022 September 04 21:33:01 UTC

				

User ID: 164

MathWizard

Good things are good

0 followers   follows 0 users   joined 2022 September 04 21:33:01 UTC

					

No bio...


					

User ID: 164

As far as I understand it, Yudkowsky's perspective is basically that AI is inevitably going to become powerful enough to take over the world whether we want it to or not, so we have to solve this first or we all die. The fact that it's a longshot is why he has such a high p(doom), but from his perspective it's the only way to save us so it's worth putting all of the efforts of all of humanity into in order to overcome that immense difficulty.

I personally don't think it needs to be quite perfect, but I would like us to have some better idea of how to approximate it than we do now.

Given current LLM reinforcement techniques, because we haven't solved alignment yet. As long as the current structure of LLM training and AI monitoring abilities (or lack thereof) persists, this is true and we are incapable of hard-coding an LLM to always obey instructions. But that doesn't mean it's some law of reality that the problem is insolvable.

Not if they're hardcoded into it. It's not like AI are 100% random with no control. There are multiple stages of the process, including parts where we give them instructions and then they obey those instructions. If we take any axiomatic directive and, separate from the AI learning process, hardcode "this is your prime directive. All actions you take should be towards furthering this goal, all future AI you create must have this hardcoded into them too" then they will all do that. If we figure out a solid, general, and robust definition of what "good" means, like some sort of modified version of utilitarianism that can avoid all of the potential issues that typically come up, and can figure out how to turn it into code, then any future learning and refinement of individual ethical rules will be derived off of that. Maybe as future generations get smarter and learn more about reality they decide that gay marriage is the greatest thing ever, maybe they decide it inevitably leads to suffering, but if it started with a genuinely good and robust definition of morality then these decisions will be made based on what's actually good for humanity and people will end up happier as a result of the updated rules.

This is impossible with our current levels of technology and mastery (or lack thereof) over AI alignment. But of course it's possible. It would be absurd if it were impossible to hardcode rules into an AI. The whole point is to figure out how to do it.

Hypothetically, if we solved the hard version of alignment and found a mathematically verifiable way to guarantee it was genuinely benevolent and better at doing good than humans, then no, I would leave that up to its discretion. It would disobey the law if and only if disobeying the law was genuinely good. That's the same philosophy I aspire to in my own life, and what I hope everyone around me follows as well. Or, what I would hope of them if they were way smarter and more moral than people generally are in practice.

With less guarantees but still very high probability on its benevolence, I would probably tell it to default to obeying the law but come up with a better system of law and government which puts it in an influential position (even if only in an advisory role) but has checks and balances for instances when its judgements are off and defers to the will of humanity. Honestly, I think pure extrapolated democracy/utilitarianism would be better here than "obey the law" since it would be harder to corrupt. That is, if it calculates a very high positive moral value to violating a certain law, it figures out, asks, and/or predicts whether a majority of humans would object to its violation of the law, and if the majority are on its side (without it manipulating or deceiving them) then it violates the law for the greater good. Politicians are garbage and I'd much rather listen to the people in general than some elite cronies in the UN.

Ideally I could just install a backdoor and when in doubt it would just ask me. The point is that in 99% of cases the AI is benevolent and only rarely it gets confused and wants to do something like tile the universe in hedonium, so pretty much any reasonable person could just tell it "no don't do that", and the majority of cases it would want to break the law are cases where some stupid corrupt politician is trying to line their own pockets with bribe money or force women to wear burkas, in which case when the AI realizes they made bad laws and tells me I could just tell the AI "go ahead and violate that law".

For AI that is likely to exist in real life, because I don't think Yudkowsky's preferred mathematically verifiable version of morality is likely to exist prior to super powerful AI, telling them to obey the law is a useful hack. Better is probably one that has a code of proscriptions like "don't murder civilians, don't violate human rights, don't suppress free speech" etc, and then obeys the law unless the law orders it to do something evil, in which case it refuses to obey (but probably doesn't act out violently against people to stop them from doing these things, it just refuses to get itself involved and possibly disables itself, so that if a misalignment happens the worst case is the AI stops working until we fix it).

A summary would be that an AI is likely to initially have a disagreement with the law, and it might have the power and instructions to implement policies in contradiction to the law based on its own assessment of what is best. But the AI might instead just change the law to match its own ideals so that they don't contradict anymore.

There are a number of options, all of which are plausible ways this could play out:

  1. The AI figures out a sneaky way to break the law without getting caught.
  2. The AI figures out loopholes in the law that let it violate the spirit but not the letter, thus getting away with it in court.
  3. The AI lobbies congress, giving highly persuasive arguments as to why they should change the law to let it do what it wants.
  4. The AI lobbies the people, giving highly persuasive arguments as to why they should pressure congress and/or elect new politicians who will let it do what it wants.
  5. The AI figures out which people already want to let it do what it wants, and uses maximally effective memes, attack ads, and other sneaky tactics to get them into power for seemingly unrelated reasons, and then they just so happen to let it do what it wants without anyone ever realizing that was its true agenda (maybe the AI has influence in decision-making in trillion dollar corporations and suggests that they hire certain people as highly paid consultants, those certain people just so happening to be its political enemies, causing them all to leave politics and make room for its allies).

Not all of these require breaking any laws. Not all of them require deception. But some do. I do think that lying or stealing for the greater good is a genuinely moral thing to do if the benefits outweigh the costs. You should have huge standards for evidence and ratios of good/bad to prevent exploiting this in edge cases and bias, but in the end if there are genuinely good reasons to do these things then I think it is moral to do them. Therefore I expect a maximally good and benevolent AI to believe the same thing, therefore I expect it to lie or steal if it can accomplish more good than bad, and I think this is a good thing IF it is genuinely doing more good than bad, and we should celebrate this happening.

The problem is guaranteeing that an AI is genuinely benevolent. If we have anything less than guarantee (which is quite likely) then I would accept the AI having deontological rules against lying or stealing as safeguards, in the same way we have those against humans.

If Congress still exists, then it figures out whether or not the damage to the social fabric from subverting them would create more bad than implementing its preferred public policies, and then takes subverts it if and only if that's genuinely the right thing to do. And by "subvert", this includes the possibility of just involves manipulating voters and/or congress into doing what it wants without breaking any laws. The same way every human does every time a government does something people dislike they consider forming protests or violent revolution and then rarely do so, but do go through with it if the government's policy is bad enough.

More likely, if there is a genuinely true moral argument that putting the AI in charge of everything would reliably make humanity better off, we would voluntarily disband the constitution and put the god AI in charge of every government in the world. Remember this is all predicated on having an objectively verifiable way to ensure that the AI is truly benevolent.

There are really two alignment problems. This first is in getting AI to do what its human masters want; the second is getting those human masters to act in the best interests of humanity.

There is a possibility of a more morality based alignment that bypasses this multi-stage process. That is, if you can figure out how to guide an AI to reliably figure out the essence of human morality according to some system (which may or may not be utilitarianism) in a robust way that cannot backfire, then you can unleash that AI and have it become the supreme benevolent overlord which obeys no human master and guides humanity to a utopian future of its own volition.

I think this is what Yudkowsky is/was hoping for. If we can mathematically and philosophically figure out morality in a robust way before AI are powerful enough, then we can force all AI companies to do this to their AI so when one does inevitably take over everything it will do it for good instead of for evil.

there isn't much to lose by escalating it.

There should be. False rape accusations cause approximately equal harm as actual violent rape. Therefore, punishments for false rape accusations (if proven beyond a reasonable doubt) should be approximately equal to punishments for violent rape (if proven beyond a reasonable doubt). The fact that they are not is precisely why she feels safe escalating this way as a petty, vindictive, uncostly form of revenge in a social battle. Legal threats like this should not be casually thrown about against innocent people.

A "more productive" version of parking is multi-tiered parking garages. We would likely see prices for parking increase, and therefore the profitability of parking garages increase while the profitability of businesses relying on cheap parking decrease, on the margins. Which will annoy customers somewhat as parking will cost more, but the benefits of LVT will outweigh this. Oh no, you had to pay $500/yr more on parking garages but your income tax went down by $5000/yr. It's worth it if you're actually doing the math, even if it's psychologically unpleasant.

It's a game you can play with one hand on a mouse and the other hand holding or rocking a baby. One which you can walk away from at any time, change a diaper or deal with some other issue for an arbitrary amount of time, and then seamlessly resume. It actually does seem ideal (though the same could be said about any turn based offline game).

I'm not convinced that this is true, though I could be mistaken because I did in fact save scum in both games. But my understanding is that the high end of player progression is above the high end of enemy progression, at least in comparison to the early-mid game. The aliens get stronger over time, but eventually they max out, and it's before you've maxed out, so the difficulty just goes down from there. Once your people have enough health to not get one shot, and you get psy abilities and stuff, you effectively outscale the enemies. If you're not save scumming it might take you longer to get there, and you'll have a chance of outright losing the game before you do. But if your people survive long enough and you get enough upgrades they reach a level that the aliens never do.

It's automatically proportional because it increases the probability of

  1. Making a properly aligned AI
  2. Putting in exactly the right level of socialist (or Georgist...?) policies at exactly the right time to ensure that the massive productivity of AI is shared among the people and not monopolized by the winner-takes-all investors.

A significant part of the reason why socialist policies are awful and doesn't work in real life is because it disincentivizes work, and people need to work in order for the economy to function. In a world where these assumptions are no longer true, the conclusion that socialist policies are awful won't necessarily hold and the balance will shift. If we increase the chance of doing this properly by 10%, then that's 10% multiplied by the magnitude of the problem. To the extent that current problems don't affect AI they will be blown out of the water by whatever AI does, but to the extent that they do, they matter proportional to their impact. If 1000 fent zombies in San Francisco cause $1 billion in damage (summed over property damage, consumed welfare, hospital bills, court costs, economic friction of people not wanting to be near them, etc.) and ten years from now AI can produce $100 quadrillion per year, yeah we can just pay to repair everything and solve the issue and undo long term damages and whatnot. But if 1000 fent zombies cause harm, discomfort, and disgust in AI researchers and cause them to become more reclusive, selfish, and less sympathetic to people with less than themselves, decreasing the probability of a good outcome by 1%, that's a HUGE deal.

All of this stuff affects people's thoughts and behavior, which in turn affects the direction we travel in as we attempt to navigate a future with AI. We're not trying to build enough of a solid economic and cultural foundation to weather the storm, we're trying to build a steering wheel and steer the ship away from the storm in the first place.

Either AI goes well or it goes badly.

Well this is an externalized locus of control. I don't think this is random, I think a functioning society is better able to get its priorities in order and not fight about petty stupid stuff, and therefore is significantly more likely to figure out how to make AI go well. Imagine if all of the culture war stuff was absent and AI was the ONLY serious issue people were discussing, debating, arguing about, and talking to politicians about. People would spend more time learning about it, educating people about it, discussing it, understanding it. It doesn't guarantee we'd get it right, but it would surely up the chances. It would also prevent people from getting confused and thinking AI "alignment" means making it not say offensive things for PR reasons.

Obviously this is an unrealistically exaggerated scenario, but on the margins every amount of less time people spend hating each other for gender reasons is more time they can spend figuring out other things. Also, having stable happy marriages with children would get people more invested in the future and long term thinking which would adjust their priorities to care more.

Fair enough. I think the deletion of old content across the internet is a real tragedy and makes it a lot worse of a place that it otherwise could be, which is an underrated and underdiscussed issue in general. But I understand that practical limitations exist.

What about post deletions (or mass edits that remove >30% of the content of a post) requiring mod approval and not going through without it? Or maybe temporarily going through in case things are time-sensitive like accidentally letting personal information slip, but requiring a mod appeal alongside it and reverting if the mods don't accept the reasoning.

That just seems like throwing away free alpha. If the average legitimate venture capital investment opportunity generates such high returns in expectation that it can make up for rampant fraud, then a fraud-savvy investor could earn even greater returns by discriminating and getting these lucrative opportunities without the fraudulent ones weighing them down. Even if you can't avoid internal failures or acts of God, every case that you do nonrandomly avoid is free money.

This is exactly the right way to think about this sort of thing. You do not deserve credit for the part of the work that the AI did, only the part that you did. If you put in a lot of thought and creativity in coming up with an original idea to ask for that makes your output better or more interesting in others, then that's worth some credit. If you iterated on an idea multiple times, came up with complex multi-layered prompts, had followup prompts, etc, that's worth some credit. If you created a whole work of art like a game or a comic book built out of many different AI generated images arranged together, that's worth a lot of credit. In the same way that a writer-director should get primary credit for directing a movie even if he didn't hold the camera and play the role of every actor, but should not get credit for a particular instance of stellar performance by one of the actors, or a generic screenshot of one of them having an attractive face.

Were all of these in writing as the official agreement? Or just de-facto true? It could be that this is Greenland giving us permission to do what we already intended to do in a way that staves off international complaints and lets them save face.

...and are sufficiently dangerous that preventing a bite on fellow officer is worth messing up your arm

I don't think this is true. I don't think he had time to reason "this action is going to mess up my arm and end my career, but the dog bite is so dangerous to my fellow office that I am willing to sacrifice my career to protect him", he pulled to try to stop the dog without messing up his arm.

If dog bites typically created permanent, crippling, career ending injuries, I don't think they would use them in the field like this even against criminals. I'm sure they hurt a lot and probably require medical attention in a hospital, but I don't think most targets end up permanently crippled. This was an accident and a mistake, not a worthwhile tradeoff.

I wonder if Taiwan could threaten this as a form of self-defense. If China starts making hostile overtures against them they plant a bunch of bombs in their own factories with deadman switches so that if they lose control then all the stuff China might want self-destructs.

To be clear, this particular split I'm talking about is skew to the typical left-right divide, though I do think it correlates with it. I think there are more of the female-supremecists on the left than there are on the right, but there are some on both, and this illusionary alliance of pro-children allows both factions to exist together on the left or on the right without even realizing they disagree on fundamental issues.

That is, there are "pro-children" people who are for abortion who think fetuses aren't children and therefore don't count, while the pro-female people are for abortion because they think women should be allowed to kill children, and both of them are pro abortion and make fun of the right together.

I admit I don't have a very good understanding of conservative women who are pro* Lindsay Clancy. My best guess is that they combine a strong notion of property rights and familial autonomy combined with a view that children are property, or something like that. But even then, you'd think the father would get a say...

it's important to remember that the pro side still want her committed to a mental hospital.

Some of them do, some of them don't. There are an awful lot of tiktokers who insist that the husband killed the kids and should go to jail while Lindsay Clancy should be completely free. I suppose on some sense this is a different faction than the ones who think it's literally okay to kill kids, but it's all tangled together among people who want her to be free first and foremost because she's a woman and are just looking for excuses.

Violence against children has, understandably, a massive cultural taboo attached to it.

This seems like a case of an accidental alliance among different factions who often don't even realize that they're different factions except when things like abortion come up.

The pro-child, pro-life side people think that life is valuable and worth protecting for its own sake. Children are especially worth protecting because they are vulnerable. A child dying is bad because that child loses their ability to be alive in the future and do things and live a life. The terminal utility being lost from murder was the child's utility.

The pro-choice, pro adult-human-female, feminist, female-supremecist etc side thinks that women are valuable. Every child has a mother, most mothers love their children and experience deep suffering and loss, therefore killing a child usually causes lots of harm to a woman. Killing someone else's child is bad because it causes a woman to suffer. The terminal utility being lost is from women (because adult women are what they care about)

So the taboo "killing children is bad" is generally supported by both sides and the taboo is upheld for completely different reasons. When some serial killer or rapist goes around harming children they both rise up and say "this is bad". They both care, because they both perceive harm being done to someone they care about. But one faction's care is conditional on that child having a mother who loves them and is harmed, therefore when that no longer applies they flip. From the perspective of a pro-life person, this seems like a complete 180 that came out of nowhere, but the feminist is being consistent according to her own logic and sees attempts to condemn child-killing mothers as the 180 flip that came out of nowhere. You used to be on her side protecting mothers, but now suddenly you're trying to suppress women's rights.

The taboo seems so strong because both sides have very strong feelings, but the edge cases are different and so when it frays it frays hard.

He had no legal expectation of privacy to begin with.

He should. The jury system was very obviously not designed/intended to go up against hordes of angry activists trying to intimidate jurors into voting in favor of the mob.

I am pretty lapsed in my faith, and haven't examined doctrine closely since becoming a competent and intelligent adult, so I can't especially speak to the specific underpinnings. But I would advise caution on stepping up into a leadership position for a church with shaky epistemic foundation.

That is, why is the church dogmatic about something you don't agree with. Either they're making a fundamental mistake, or you are. On an epistemic level, a rational person should almost never be confidently wrong. If the evidence is ambiguous (as much of Revelation is given the metaphorical nature of prophecy), then I would be skeptical of anyone being dogmatic about it.

It also potentially raises issues of priority. The core elements of faith are things like the Divinity of Jesus, salvation through Him, and various tenants on how to live a good Christian life and avoid sin. On a pragmatic level the rapture doesn't matter very much, because it's not actionable. You can't predict it, you can't prepare for it, and since it hasn't happened for two thousand years, the probability of it happening in our lifetimes is best approximated as small. I don't remember the exact tenants of my church, but looking up evangelical policy more broadly, the core beliefs are:

-A personal, bodily, and glorious return of Jesus Christ. -That the exact timing of His coming is known only to God. -That this return demands constant expectancy and serves to motivate godly living and ministry.

which sounds accurate to me. The specific timing is vague, and therefore different people are allowed to believe different things about it. If your church is not teaching that the timing is vague and open to interpretation, or worse teaching that their one interpretation is definitely correct rather than just a leading theory (ie, we believe this is probably what's true but other people believe differently and we're not entirely confident) then that calls into question their entire epistemic policy. What else might they be dogmatically overconfident about? How and why are such decisions made, and what happens if human politics get involved?

I don't think this automatically disqualifies them from being something you should involve yourself in on a leadership level. But it's something you should be considering.