Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Sopranos illustrating Melfi Ends Tony’s TherapyScene still / The Sopranos
Archive noteNo clip in the collectionThe scene analysis remains available in full.

Melfi Ends Tony’s Therapy

After reading research arguing that sociopaths can use talk therapy to become better manipulators, Melfi recognizes that Tony has been turning treatment into strategic support and terminates their eight-year relationship.

An apparently interpretive oversight process may itself furnish a strategic actor with better self-models, vocabulary, and manipulation tools. Melfi's final safeguard is not another conversation but withdrawal of access.

Tony is a human patient, the clinical claim is debated, and Melfi's abrupt termination raises its own ethical questions; this is an analogy about oversight becoming capability support.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Scheming
  2. Corrigibility
  3. Situational Awareness
  4. Sycophancy