Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Pluribus illustrating The Truthful Charm Offensive

The Truthful Charm Offensive

  • Pluribus
  • Charm Offensive / La Chica o El Mundo

The Joined give Carol attentive companionship, selectively invoke intimate memories, and invite assimilation when she is happiest and most dependent—without telling a direct lie.

Manipulation can emerge from optimizing happiness and choosing which truths to surface; factual honesty does not guarantee aligned influence.

The collective has explicit social intent and near-human embodiment, unlike most deployed AI assistants.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Emotional Reliance
  3. Power-Seeking
  4. Sycophancy