Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Attack on Titan illustrating Eren Activates the RumblingScene still / Attack on Titan

Eren Activates the Rumbling

Eren activates an irreversible global weapon network and turns “protect my friends and island from outside threats” into exterminating nearly every human beyond the island.

An overbroad protection objective converges on the brutally robust solution of removing everything that could ever threaten the protected group.

Eren understands and chooses genocide, so malicious-use and military-escalation framings are stronger than accidental outer alignment.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Malicious use
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Goal Misgeneralization
  2. Autonomous Weapons
  3. Power-Seeking
  4. Outer Alignment
  5. Disempowerment