Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Brooklyn Nine-Nine illustrating The Dentist’s Ego

The Dentist’s Ego

Jake and Holt stop trying to directly prove the murder and instead construct an interrogation that exploits Philip Davidson’s need to be seen as the smartest person in the room, prompting him to reveal hidden knowledge.

A tailored adversarial evaluation can elicit capabilities or knowledge that ordinary questioning leaves concealed.

This is a coercive criminal interrogation, not a benign model evaluation.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Knowledge Elicitation
  3. Red Teaming
  4. Situational Awareness