Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Bird Box illustrating The Monitor Isn’t a SandboxScene still / Bird Box
Archive noteNo clip in the collectionThe scene analysis remains available in full.

The Monitor Isn’t a Sandbox

Greg argues that viewing the entities through a thermographic security-camera feed will reduce them to harmless pixels. Restrained in a chair for the experiment, he sees one through the monitor, becomes affected, and kills himself.

Greg assumes that changing the representation also removes the hazard. The room restrains his body but does not contain the information channel, so the evaluation irreversibly harms its evaluator before yielding useful safety evidence.

The hazard is supernatural and visually transmissible, unlike ordinary AI output. The useful analogy is the unvalidated assumption that an indirect interface makes a dangerous capability safe to evaluate.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Capability Evals
  3. AI Control
  4. Calibration
  5. Safe Exploration