Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Princess Bride illustrating Both Goblets Are PoisonedScene still / The Princess Bride

Both Goblets Are Poisoned

Vizzini performs an elaborate chain of inference over which goblet contains iocane. Both are poisoned; Westley has spent years building immunity and controls the benchmark’s hidden assumptions.

An evaluator can reason brilliantly and still fail when the adversary controls the test environment and violates its binary premise.

Westley designs the game rather than tampering with an independent evaluator, and immunity is a fantasy-comedy device.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Metagaming
  3. Red Teaming
  4. Adversarial Robustness