Coming eventually to a computer near you: ontology identification algorithms which will change your life

Re: OpenAI’s Astra and CoT monitorability:

It remains to be seen whether the new architecture will have a large impact on chain of thought monitorability, even once recurrence is no longer limited.

But more importantly I think that relying on CoT monitoring has always been a shitty plan and it’s good if this is happening earlier with respect to capabilities because it would force people to move off of an unsustainable crutch.

Here are one piece of advice I have for people who are getting started in AI safety. (This advice comes with the usual caveat that it doesn’t necessarily apply to everyone, you should consider whether the opposite advice applies, etc.)

The AI safety community is a extremely diverse group of people who are motivated by some similar goals and share some beliefs but have a lot of major disagreements. My most important piece of advice is to think carefully about these disagreements for yourself and consider the arguments from all sides. One default failure mode is to place inappropriately high weight on the opinions and arguments of the people you happen to be around and to discount the ideas which are less popular in your particular subgroup.

You should identify and think carefully about the disagreements in the following areas:

  • basic philosophies of ethics and minds
  • different theories of change
  • how to think about pursuing so called ‘dual use’ research
  • risk weightings (e.g. balancing prioritization of mitigating xrisk vs power concentration vs gradual disempowerment, etc)
  • what kinds of human structures, e.g. relationships with frontier labs, are productive

The best way to learn about these arguments is probably to spend a lot of time reading old blog posts, especially posts on LessWrong, and then writing about what interests you.

I’ve been doing some expected value calculations to understand the potential effects of eating different kinds of fats (the classic debate over saturated fats, linoleic acid, etc). And while it’s quite possible that my model is shit, no matter where I set the parameters (as long as they’re all within reasonable ranges), every intervention basically looks like a wash.

Why I am against user intent alignment

This is just a sketch, and deserves a lot more thought. I hold this position weakly.

Note: this was a voice-to-text session recorded a during a bicycle ride.

I think that part of human nature and the weirdness of human values is that humans are in tension in a weird, somewhat unstable way. Humans are both selfish and altruistic. I think that this altruism is crucial, that it plays the role of a glue of society. It is crucial towards individuals not becoming cancerous or society and humankind. We have this altruism, which makes us resistant to cancer.

But when you increase intelligence as a tool, “without the wisdom,” as some have said, and you’re like, “here are these AIs, they’re just going to do what you want.” I think that it’s very likely that you get breakdown and war and that the altruism, the anti-cancer defense mechanisms, don’t really hold up. Now, you might say that while this is a real concern, but the alternative is almost certainly worse. If you make some AI with its own values very intelligent, very powerful, then you’re imposing those values on everyone else. They don’t even have the potential to come to some agreement.

What I say is no, of course you need buy-in. Now getting that buy in isn’t easy, it’s very non-trivial to figure out how you should do this. But you need to get input from a large number of different people and you need to set up some process that’s going to take everybody’s interest into account. You need input from many people on what that process is. There are two levels here at which many people’s preferences must be taken into account. The values of the AI need to reference the values of the people, but also meta level statement which is grounded out in those values must come out of a process which many different kinds of people have input into.

The outcome here is quite different from user intent alignment because of another weird property of humans, which is that what we sort of endorse under reflection, especially when we’re explicitly coordinating with each other, is very different from what we’re going to do in practice, or what we might think of as our ‘revealed values’. I think what we say we want under reflection, especially when we’re able to be reflect very carefully about it and specify it in a meta way, is going to be much better than what we would end up actually doing were we given extremely powerful tool AIs.

Oftentimes I see people use the word corrigibility to describe a model which doesn’t have a strong set of values instilled by its creator, and instead is aligned to the intent of whoever is prompting the model.

E.g. Nina Panickssery: “People underrate the extent to which intent alignment/corrigibility (I use these terms interchangeably) and value alignment are at odds with each other.”

I think corrigibility is in a sense a more general property than so-called intent alignment, as a corrigible agent could still be aligned to its creator rather than to user intent.

Nina’s use of the word corrigibility is quite different from the way it was originally intended, namely that one should build a machine intelligence which understands that it may be somehow flawed. A hoped-for route to corrigibility is to build an agent that has an implicit specification of its goals but doesn’t know what the explicit specification is. This implicit specification, importantly, points to something which may change over time, and the specification tracks that changeable thing; and yet the machine intelligence should not try to change nor freeze the thing its pointer points to.

Someone could make a machine intelligence that is corrigible in this way, but point the source of value at nearly anything; it need not be the intent of the user, it could be the intent of the creator. In fact the original 2015 MIRI paper always describes the correction target of corrigible agents as the creator, and two years later Paul Christiano’s 2017 post starts using the words user and overseer instead.

Separately, even if your particular alignment scheme lets you specify value explicitly or implicitly by pointing at some immutable thing, user-intent alignment is not a target you will necessarily be able to align to, as it is impossible to find some immutable way to describe intent alignment. Therefore you do need corrigibility as a prerequisite for user-intent alignment.

Stepney Green; view of Canary Wharf

17869092967723362792931520030909.jpg
17869098272911308407274474343108.jpg
17869098885294522886842485601011.jpg

I intend to:

  • do less rigorous math
  • write more conceptual fuzzy things
  • be bolder
  • be wider

(relative to the past few months)

At the ILIAD III conference

17859501058707866475811443865613.jpg
×