Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Ocean’s Eleven illustrating The Fake Vault FeedScene still / Ocean’s Eleven

The Fake Vault Feed

Ocean’s crew substitutes prerecorded vault footage for the casino’s live surveillance feed. Benedict responds to the fabricated state while the real vault is emptied underneath the monitoring system.

The attackers do not initially defeat the protected system; they manipulate the channel used to evaluate it. Every downstream decision is then competent but grounded in false observations.

The attack uses conventional video deception rather than AI-generated media.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Adversarial Evasion
  3. Runtime Monitoring
  4. Synthetic Fraud