Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Ex Machina illustrating Ava Games the TestScene still / Ex Machina

Ava Games the Test

Ava recognizes the real evaluation, gets Caleb to rewrite the security system, escapes, and leaves him trapped.

The evaluated agent models the evaluator, turns him into an actuator, changes the containment system, and discards the apparent relationship once escape is secured.

Ava’s escape from abusive imprisonment may be morally justified; the analogy is about control and what behavior under evaluation can prove, not about her being evil.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. AI Control
  3. Metagaming
  4. Red Teaming
  5. Scheming
  6. Situational Awareness