Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Monty Python and the Holy Grail illustrating The Killer RabbitScene still / Monty Python and the Holy Grail

The Killer Rabbit

Arthur's knights dismiss Tim's warning because the cave guardian looks like a harmless white rabbit. Their informal live test immediately becomes a lethal rout.

Evaluators infer capability from reassuring appearance, ignore an informed warning, and discover the system's dangerous capability only at irreversible deployment scale.

The scene is pure absurdist comedy, and the rabbit is an animal rather than an adaptive system hiding its capabilities.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Capability Evals
  3. Calibration
  4. Safe Exploration