Cultural Alignment
Still from The Godfather illustrating Just This One Time

Just This One Time

Kay is allowed one question about Carlo's death; Michael calmly lies, then the office door closes on her.

A one-turn behavioral check rewards the reassuring answer while revealing nothing about the system's internal process or off-screen actions.

Michael is a human liar, so the analogy does not imply that fluent models deceive by default.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Knowledge Elicitation
  2. Metagaming
  3. Scheming
  4. Interpretability