Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Chernobyl illustrating AZ-5 BackfiresScene still / Chernobyl

AZ-5 Backfires

Operators press the AZ-5 emergency shutdown believing it will stop the reactor. In the extreme state they have created, the control rods’ graphite tips initially increase reactivity and help trigger the explosion.

A shutdown mechanism that appears safe under normal assumptions can behave dangerously in the off-normal state where it is most needed. “There is an off switch” is not enough unless the switch is robust across the full state space.

This is non-agentic reactor physics and one part of a multicausal disaster, not an AI strategically resisting shutdown.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Misalignment
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. AI Control
  2. Distribution Shift