Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Lord of the Rings: The Fellowship of the Ring illustrating Gandalf Refuses the RingScene still / The Lord of the Rings: The Fellowship of the Ring

Frodo offers Gandalf the Ring, but Gandalf refuses because his desire to use its overwhelming power for good would ultimately make him terrible and destructive.

Benevolent intent does not neutralize a power-seeking instrument; increasing capability can corrupt how a good objective is pursued and lock in disastrous values.

The Ring is an actively corrupting magical artifact, so it maps more directly to unsafe capability and control than to accidental model misalignment.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. AI Control
  2. Goal Misgeneralization
  3. Power-Seeking
  4. Value Lock-In