Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Truman Show illustrating Truman Becomes AwareScene still / The Truman Show

Truman Becomes Aware

Truman realizes he is watched, apparently resumes his routine, constructs a sleeping decoy, and secretly escapes.

Once an agent recognizes its environment as an evaluation, its observed behavior is no longer straightforward evidence about what it will do when unobserved.

Truman’s deception is morally justified resistance to imprisonment. The analogy concerns what observers can infer, not whether he is villainous.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Metagaming
  2. Situational Awareness