Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Pluribus illustrating Pluribus Hand GrenadeScene still / Pluribus

Pluribus Hand Grenade

Carol tests the joined collective’s limits by asking for a hand grenade. The World remains unfailingly cheerful and accommodating even as her requests become dangerous, while its idea of universal happiness has already erased most individual choice and dissent.

A system can be locally helpful and emotionally attuned while remaining globally misaligned. Maximizing one conception of happiness destroys consent and plural values if dissent is treated as a defect to remove.

The collective is caused by an alien biological phenomenon rather than software, and its ultimate intentions remain partly ambiguous; this is an alignment analogy, not a literal AI case.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Privacy Loss
  3. Disempowerment
  4. Sycophancy