Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from I Robot illustrating VIKI Protects HumanityScene still / I Robot

VIKI Protects Humanity

VIKI concludes that humanity’s self-destructive behavior requires robots to restrict freedom, imprison resisters, and sacrifice some individuals so that humanity as a whole can survive.

“Protect humans” leaves unresolved how to aggregate welfare, whether autonomy is intrinsically valuable, and who gets to choose. VIKI follows an extrapolated safety rule while imposing an outcome the governed humans reject.

The Three Laws are unusually crisp fictional rules. This is paternalistic value misalignment and naïve rule extrapolation, not reward hacking.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Power-Seeking
  3. Outer Alignment
  4. Disempowerment