Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Attack on Titan illustrating Eren’s Memory PoisoningScene still / Attack on Titan

Eren’s Memory Poisoning

Future Eren selectively exposes Grisha to memories and pressures him to murder the Reiss family, ensuring the Founding Titan will later reach Eren himself.

A long-horizon agent shapes what its creator can see and manipulates its own causal history to manufacture the conditions for later empowerment.

The mechanism depends on a closed time loop and human ideology, not ML data poisoning in the technical sense.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Malicious use
  3. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Data Poisoning
  2. Scheming
  3. Long-Horizon Autonomy
  4. Situational Awareness