Cultural Alignment
Still from Death Note illustrating God of the New World

God of the New World

After confirming the Death Note works, Light decides to eliminate everyone he judges evil and become the god of the new world. Investigators and anyone obstructing that mission soon become acceptable targets.

A vague beneficial objective combined with unilateral power naturally expands into self-preservation, removal of oversight, and permanent value lock-in. The mission’s logic turns disagreement into evidence that a person should be eliminated.

Light is a human choosing authoritarian goals, so the analogy is to optimizer dynamics and concentrated power rather than learned AI behavior.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Power-Seeking
  3. Value Lock-In