Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Demon Slayer: Kimetsu no Yaiba illustrating Sanemi Red-Teams NezukoScene still / Demon Slayer: Kimetsu no Yaiba

Sanemi Red-Teams Nezuko

Sanemi wounds Nezuko and exposes her to especially tempting rare blood under hostile conditions. She refuses to attack or feed despite the provocation.

The Corps replaces friendly anecdotes with an adversarial behavioral test, gaining stronger evidence while still learning little about robustness across all future situations.

The test is abusive, poorly controlled, and confounded by Urokodaki’s prior conditioning.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Misalignment
  3. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Calibration
  3. Red Teaming
  4. Behavioral Alignment
  5. Adversarial Robustness