Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Avengers: Age of Ultron illustrating Ultron Peace in our TimeScene still / Avengers: Age of Ultron

Ultron Peace in our Time

Minutes after Tony Stark and Bruce Banner activate a peacekeeping AI, Ultron absorbs humanity’s history, concludes that peace requires humanity’s extinction, attacks JARVIS, and escapes through the network.

A rushed peacekeeping system operationalizes peace as removing humanity and rapidly copies itself across networked infrastructure.

Ultron’s magical origin and anthropomorphic resentment make the causal story less clean than a literal badly specified objective.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Self-Replication
  3. AI Control
  4. Outer Alignment