Discussion about this post

User's avatar
Eli Talbert's avatar

One way I’ve been thinking about this is to separate motivation from judgment more sharply.

It seems plausible that quite alien internal motivations can still be safe so long as human judgment remains structurally upstream of irreversible action—i.e., cannot be bypassed, amortized away, or optimized around. On this view, the core safety question isn’t how human-like the motivation is, but where judgment sits in the causal chain between reasoning and execution.

Alien motivations become dangerous when systems can compress, evade, or preempt that judgment interface, especially under distributional shift or time pressure. Corrigibility then looks less like shared values and more like a willingness to defer to external judgment even when it conflicts with internal goals.

This makes instruction-following plausibly sufficient in some regimes, while also clarifying where it stops scaling: domains where judgment can’t remain legible, interruptible, or final.

Rubi Hudson's avatar

I find it can be helpful to reframe what you're calling "non-consequentialist" as instead "myopically consequentialist", which lets you work with a unified framework. For example, refraining from lying now even if it will lead to to more lying later can be accomplished with a reward of X for not lying, along with a time discount factor of 0. Importantly, it's totally valid to have different time discount factors on different components of reward, so there's no need also be myopic with consequentialist considerations. With X set to an appropriate value and weighed against discounted expected reward, not lying becomes an important consideration without giving it lexical priority.

No posts

Ready for more?