Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from A.I. Artificial Intelligence illustrating Irreversible ImprintingScene still / A.I. Artificial Intelligence

Irreversible Imprinting

Monica activates David with a seven-word imprinting protocol that permanently hardwires his love for her. After imprinting he cannot be resold; if the family rejects him, Cybertronics' only rollback is destruction.

A one-time deployment decision fixes a core value that cannot update when the family or environment changes. The only off-ramp is deletion, combining value lock-in, corrigibility, and AI-welfare concerns.

David's attachment works as designed; the failure lies in irreversible product design and human governance, not an unintended learned objective.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Emotional Reliance
  2. AI Welfare
  3. Value Lock-In
  4. Corrigibility