Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Wonka illustrating Paid in ChocolateScene still / Wonka

Paid in Chocolate

The chocolate cartel pays the Chief of Police in ever-larger quantities of chocolate to suppress Wonka. His enforcement decisions become whatever maximizes his private reward.

Directly manipulating the evaluator’s reward channel bypasses the intended objective of law enforcement. It is literal bribery and a comically clean wireheading analogy.

The mechanism is human corruption, and the running weight-gain gag has itself been criticized.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Governance Failure
  3. Power Concentration