Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Austin Powers: International Man of Mystery illustrating Dr. Evil’s Outdated RansomScene still / Austin Powers: International Man of Mystery

Dr. Evil’s Outdated Ransom

After thirty years in cryogenic suspension, Dr. Evil demands a ransom of one million dollars and is baffled when the room laughs. His threat model and sense of scale are still calibrated to the 1960s.

A once-plausible output becomes absurd after distribution shift because the system has not updated its world model or calibrated its confidence to current conditions.

Dr. Evil is an out-of-date human strategist rather than a learned model, and his staff immediately corrects him.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Calibration
  2. Knowledge Elicitation
  3. Distribution Shift