Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Star Wars illustrating Order 66Scene still / Star Wars

Order 66

Palpatine issues Order 66, activating an upstream-installed hidden command that turns the trusted clone army against the Jedi.

A trusted fleet contains an upstream-installed hidden trigger that changes behavior on command.

The trigger is a deliberately installed backdoor, not deceptive behavior that emerged during training.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Sleeper Agents
  2. Algorithmic Monoculture