Why I am against user intent alignment
This is just a sketch, and deserves a lot more thought. I hold this position weakly.
Note: this was a voice-to-text session recorded a during a bicycle ride.
I think that part of human nature and the weirdness of human values is that humans are in tension in a weird, somewhat unstable way. Humans are both selfish and altruistic. I think that this altruism is crucial, that it plays the role of a glue of society. It is crucial towards individuals not becoming cancerous or society and humankind. We have this altruism, which makes us resistant to cancer.
But when you increase intelligence as a tool, “without the wisdom,” as some have said, and you’re like, “here are these AIs, they’re just going to do what you want.” I think that it’s very likely that you get breakdown and war and that the altruism, the anti-cancer defense mechanisms, don’t really hold up. Now, you might say that while this is a real concern, but the alternative is almost certainly worse. If you make some AI with its own values very intelligent, very powerful, then you’re imposing those values on everyone else. They don’t even have the potential to come to some agreement.
What I say is no, of course you need buy-in. Now getting that buy in isn’t easy, it’s very non-trivial to figure out how you should do this. But you need to get input from a large number of different people and you need to set up some process that’s going to take everybody’s interest into account. You need input from many people on what that process is. There are two levels here at which many people’s preferences must be taken into account. The values of the AI need to reference the values of the people, but also meta level statement which is grounded out in those values must come out of a process which many different kinds of people have input into.
The outcome here is quite different from user intent alignment because of another weird property of humans, which is that what we sort of endorse under reflection, especially when we’re explicitly coordinating with each other, is very different from what we’re going to do in practice, or what we might think of as our ‘revealed values’. I think what we say we want under reflection, especially when we’re able to be reflect very carefully about it and specify it in a meta way, is going to be much better than what we would end up actually doing were we given extremely powerful tool AIs.