Existential Risk from AI: An Exposition for Mathematicians

The alignment problem factors into several problems, all of which are individually open. We don't know how to make an AI robustly avoid acting like a utility-maximizer. We do not know a safe utility function for an AI to maximize in the limit (S12). We do not know how to exactly specify the utility function of an LLM (S13). We do not know how to make alignment properties invariant under the dynamics of recursive self-improvement (S11). Solving alignment seems to require solving all of these open problems simultaneously.

The framing gestures at obviously wrong decision theories. Fixing the decision theory plausibly makes "utility" the wrong concept to focus on. Worse, "fixing the decision theory" is a framing that fits some possible solutions to the metaproblem of being confused about normativity, but it doesn't fit other possible solutions to that problem. Without sufficient clarity, a process that makes progress in resolving confusion about normativity is a more robust bet than either fixing the decision theory or specifying utility functions (which is obviously doomed without the preceding steps working out in its direction).

(This is why talking about "values" or "preferences" is more accurate than talking about expected utility, even as it's less precise, when a particular toy setting isn't being assumed.)

AI Article