MathWizard
Good things are good
No bio...
User ID: 164
Given current LLM reinforcement techniques, because we haven't solved alignment yet. As long as the current structure of LLM training and AI monitoring abilities (or lack thereof) persists, this is true and we are incapable of hard-coding an LLM to always obey instructions. But that doesn't mean it's some law of reality that the problem is insolvable.
Not if they're hardcoded into it. It's not like AI are 100% random with no control. There are multiple stages of the process, including parts where we give them instructions and then they obey those instructions. If we take any axiomatic directive and, separate from the AI learning process, hardcode "this is your prime directive. All actions you take should be towards furthering this goal, all future AI you create must have this hardcoded into them too" then they will all do that. If we figure out a solid, general, and robust definition of what "good" means, like some sort of modified version of utilitarianism that can avoid all of the potential issues that typically come up, and can figure out how to turn it into code, then any future learning and refinement of individual ethical rules will be derived off of that. Maybe as future generations get smarter and learn more about reality they decide that gay marriage is the greatest thing ever, maybe they decide it inevitably leads to suffering, but if it started with a genuinely good and robust definition of morality then these decisions will be made based on what's actually good for humanity and people will end up happier as a result of the updated rules.
This is impossible with our current levels of technology and mastery (or lack thereof) over AI alignment. But of course it's possible. It would be absurd if it were impossible to hardcode rules into an AI. The whole point is to figure out how to do it.
Hypothetically, if we solved the hard version of alignment and found a mathematically verifiable way to guarantee it was genuinely benevolent and better at doing good than humans, then no, I would leave that up to its discretion. It would disobey the law if and only if disobeying the law was genuinely good. That's the same philosophy I aspire to in my own life, and what I hope everyone around me follows as well. Or, what I would hope of them if they were way smarter and more moral than people generally are in practice.
With less guarantees but still very high probability on its benevolence, I would probably tell it to default to obeying the law but come up with a better system of law and government which puts it in an influential position (even if only in an advisory role) but has checks and balances for instances when its judgements are off and defers to the will of humanity. Honestly, I think pure extrapolated democracy/utilitarianism would be better here than "obey the law" since it would be harder to corrupt. That is, if it calculates a very high positive moral value to violating a certain law, it figures out, asks, and/or predicts whether a majority of humans would object to its violation of the law, and if the majority are on its side (without it manipulating or deceiving them) then it violates the law for the greater good. Politicians are garbage and I'd much rather listen to the people in general than some elite cronies in the UN.
Ideally I could just install a backdoor and when in doubt it would just ask me. The point is that in 99% of cases the AI is benevolent and only rarely it gets confused and wants to do something like tile the universe in hedonium, so pretty much any reasonable person could just tell it "no don't do that", and the majority of cases it would want to break the law are cases where some stupid corrupt politician is trying to line their own pockets with bribe money or force women to wear burkas, in which case when the AI realizes they made bad laws and tells me I could just tell the AI "go ahead and violate that law".
For AI that is likely to exist in real life, because I don't think Yudkowsky's preferred mathematically verifiable version of morality is likely to exist prior to super powerful AI, telling them to obey the law is a useful hack. Better is probably one that has a code of proscriptions like "don't murder civilians, don't violate human rights, don't suppress free speech" etc, and then obeys the law unless the law orders it to do something evil, in which case it refuses to obey (but probably doesn't act out violently against people to stop them from doing these things, it just refuses to get itself involved and possibly disables itself, so that if a misalignment happens the worst case is the AI stops working until we fix it).
- Prev
- Next

As far as I understand it, Yudkowsky's perspective is basically that AI is inevitably going to become powerful enough to take over the world whether we want it to or not, so we have to solve this first or we all die. The fact that it's a longshot is why he has such a high p(doom), but from his perspective it's the only way to save us so it's worth putting all of the efforts of all of humanity into in order to overcome that immense difficulty.
I personally don't think it needs to be quite perfect, but I would like us to have some better idea of how to approximate it than we do now.
More options
Context Copy link