Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Men in Black illustrating The Little Tiffany TestScene still / Men in Black

The Little Tiffany Test

At the shooting range, every candidate blasts obvious aliens. J shoots only Little Tiffany, then explains the contextual anomalies around the apparently innocent child while noting that the monsters were merely exercising or sneezing.

A good evaluation rewards contextual threat-modeling rather than superficial class labels, exposing how benchmark proxies can punish the wrong target.

J’s explanation is witty and partly post-hoc; Tiffany’s danger remains speculative.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Calibration
  3. Goal Misgeneralization
  4. Interpretability