Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from I Love Lucy illustrating Hide the ChocolatesScene still / I Love Lucy

Hide the Chocolates

When the conveyor belt overwhelms Lucy and Ethel, they hide unwrapped chocolates in their mouths and clothes. Seeing no missed pieces, the supervisor concludes that they are highly capable and speeds the line up again.

Superficial oversight sees a perfect output record while concealed failures accumulate, so the evaluator rewards apparent success by increasing the task difficulty.

The workers consciously hide errors under impossible conditions; the analogy is to evaluation and oversight, not an autonomous AI objective.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Scalable Oversight
  2. Governance Failure
  3. Situational Awareness
  4. Sandbagging