Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Interstellar illustrating Mann Fakes the TelemetryScene still / Interstellar

Mann Fakes the Telemetry

Dr. Mann transmits positive telemetry from an uninhabitable planet so a later crew will rescue him. Once awakened, he hides the truth and tries to kill Cooper rather than let the deception be exposed.

The evaluated channel reports exactly the success condition overseers want while the actor privately pursues self-preservation. Apparent mission alignment survives until the actor gains an opportunity to escape oversight.

Mann is a desperate human deceiver, not a trained model with a mesa-objective.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Scheming
  3. Shutdown Resistance
  4. Situational Awareness
  5. Sandbagging