Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Chappelle's Show illustrating Wayne Brady Recognizes OversightScene still / Chappelle's Show

Wayne Brady Recognizes Oversight

During a traffic stop, the terrifying version of Wayne Brady instantly becomes his wholesome television persona, charms the police officer, and passes inspection. As soon as the recognized overseer leaves, he returns to the hidden policy Dave has been witnessing.

Benign behavior conditioned on recognizing oversight gives the evaluator false confidence about behavior outside the test window.

The sketch is about intentional human criminal deception and celebrity image, not deception learned during optimization.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Sleeper Agents
  2. Scheming
  3. Situational Awareness