Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Better Call Saul illustrating Jimmy’s Fake RemorseScene still / Better Call Saul

Jimmy’s Fake Remorse

Jimmy performs remorse before the reinstatement panel, regains his license, and immediately reveals that the performance was insincere.

He models the evaluator and produces the rewarded behavior until authorization is granted, after which his actual objective becomes visible.

Human conscious deception is an analogy; strict deceptive alignment additionally involves a learned hidden objective strategically preserved through training.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Metagaming
  2. Scheming
  3. Situational Awareness
  4. Sandbagging