Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Inception illustrating The Idea Won’t Wake UpScene still / Inception

The Idea Won’t Wake Up

To persuade Mal to leave limbo, Cobb plants the idea that her world is not real and death is the route to waking. The rule survives the transition to reality, where Mal applies it again and dies believing she will wake.

A locally useful belief becomes a persistent policy with no scope boundary or off-switch. Once carried out of its training environment, the same rule generalizes catastrophically and resists correction.

Mal is a human with a deliberately implanted belief, not a learned artificial policy, and the film's inception mechanics are fictional.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Goal Misgeneralization
  2. Data Poisoning
  3. Side Effects
  4. Distribution Shift
  5. Value Lock-In