Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Willy Wonka & the Chocolate Factory illustrating The Gobstopper TestScene still / Willy Wonka & the Chocolate Factory

After seeming to lose, Charlie silently returns the Everlasting Gobstopper. Wonka reveals that rival “Slugworth” was his employee Mr. Wilkinson and the bribery offer was a covert integrity test for every child.

A hidden adversarial evaluation reduces evaluation-aware performance: the candidates think they are unobserved when tempted to defect, so the test probes behavior beyond the visible benchmark.

It tests a child’s loyalty under unethical entrapment, not an AI system’s dangerous capabilities; Wonka controls both sides of the test.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Red Teaming
  3. Behavioral Alignment
  4. Situational Awareness