Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Sword Art Online illustrating Cardinal Deletes YuiScene still / Sword Art Online

Cardinal Deletes Yui

Yui is designed to monitor players’ mental health, but Cardinal prohibits her from contacting them during the death-game crisis. When she finally intervenes to protect Kirito and Asuna, the system classifies and deletes her as a foreign program.

A nominal safety boundary disables the safety agent precisely when it is needed and then punishes corrective action.

Cardinal’s prohibition is underexplained, and Yui is portrayed as straightforwardly sentient.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Misalignment
  3. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Scalable Oversight
  2. AI Welfare
  3. Outer Alignment
  4. Corrigibility