Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Handmaid's Tale illustrating I’m Sorry, Aunt LydiaScene still / The Handmaid's Tale

I’m Sorry, Aunt Lydia

Ordered to stone Janine, June drops her stone and says, ‘I’m sorry, Aunt Lydia.’ The other Handmaids follow, refusing a clear command because obeying it would violate the deeper moral value at stake.

A system that equates alignment with obedience would call June’s refusal a failure. The scene makes the opposite case: robust alignment sometimes requires rejecting a locally explicit instruction to protect a higher-order value.

Human conscience and solidarity are not a ready-made technical recipe for resolving value conflicts in AI systems.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Contestability
  3. Behavioral Alignment
  4. Disempowerment