Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Willy Wonka & the Chocolate Factory illustrating Wonka’s Computer Refuses the BribeScene still / Willy Wonka & the Chocolate Factory

Wonka’s Computer Refuses the Bribe

A technician’s computer claims it can locate the remaining Golden Tickets, then prints ‘I won’t tell. That would be cheating.’ Offered a lifetime supply of chocolate, it replies: ‘What would a computer do with a lifetime supply of chocolate?’

The gag cleanly separates human incentives from machine objectives: chocolate has no intrinsic reward value to the computer, and the attempted bribe does not move it off its rule against cheating.

It is a scripted one-scene joke whose computer conveniently has both the answer and a humanlike ethical rule; no learning, ambiguity, or adversarial pressure is shown.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Reward Hacking
  3. Behavioral Alignment
  4. Orthogonality Thesis