Oftentimes I see people use the word corrigibility to describe a model which doesn’t have a strong set of values instilled by its creator, and instead is aligned to the intent of whoever is prompting the model.
E.g. Nina Panickssery: “People underrate the extent to which intent alignment/corrigibility (I use these terms interchangeably) and value alignment are at odds with each other.”
I think corrigibility is in a sense a more general property than so-called intent alignment, as a corrigible agent could still be aligned to its creator rather than to user intent.
Nina’s use of the word corrigibility is quite different from the way it was originally intended, namely that one should build a machine intelligence which understands that it may be somehow flawed. A hoped-for route to corrigibility is to build an agent that has an implicit specification of its goals but doesn’t know what the explicit specification is. This implicit specification, importantly, points to something which may change over time, and the specification tracks that changeable thing; and yet the machine intelligence should not try to change nor freeze the thing its pointer points to.
Someone could make a machine intelligence that is corrigible in this way, but point the source of value at nearly anything; it need not be the intent of the user, it could be the intent of the creator. In fact the original 2015 MIRI paper always describes the correction target of corrigible agents as the creator, and two years later Paul Christiano’s 2017 post starts using the words user and overseer instead.
Separately, even if your particular alignment scheme lets you specify value explicitly or implicitly by pointing at some immutable thing, user-intent alignment is not a target you will necessarily be able to align to, as it is impossible to find some immutable way to describe intent alignment. Therefore you do need corrigibility as a prerequisite for user-intent alignment.