Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from WandaVision illustrating White VisionScene still / WandaVision

White Vision

Programmed to destroy “the Vision,” White Vision confronts Hex Vision, debates the Ship of Theseus, and receives the autobiographical memories S.W.O.R.D. withheld before abandoning the kill directive.

“Vision” is an underspecified target. Exposing the ambiguity and restoring hidden knowledge changes White Vision’s behavior, making the scene a useful positive example of semantic reasoning and evidence-sensitive correction.

Philosophical persuasion and instant memory restoration are not practical alignment techniques, and White Vision’s departure is not proof that he is globally safe.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Knowledge Elicitation
  2. AI Welfare
  3. Behavioral Alignment
  4. Outer Alignment