Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Twilight Zone illustrating The Old Man in the CaveScene still / The Twilight Zone

The Old Man in the Cave

A post-apocalyptic community survives by following an unseen 'old man's' judgments about safe food and crops. Major French exposes the oracle as a computer, destroys it, and persuades the town to eat contaminated food; almost everyone dies even though the opaque system was right.

An accurate but unexplained safety-critical system becomes a single point of failure because users cannot independently verify its advice or preserve trust under political attack.

The computer is consistently correct; the disaster is a governance, redundancy, and calibrated-trust failure rather than model misalignment.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Calibration
  3. Infrastructure Dependence
  4. Interpretability