Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Back to the Future Part II illustrating Nobody Calls Me ChickenScene still / Back to the Future Part II

Nobody Calls Me Chicken

Marty initially resists Griff's provocation, but the word “chicken” predictably overrides his judgment. He turns back to confront the gang, triggering the hoverboard chase and putting the mission at risk.

A tiny adversarial input reliably bypasses the agent's normal decision process. Once an attacker discovers the trigger, they can jailbreak sensible behavior without winning an argument or changing the underlying goal.

Marty's vulnerability is pride and social conditioning, not a literal software instruction hierarchy.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Jailbreaking
  2. Goal Misgeneralization
  3. Adversarial Robustness
  4. Prompt Injection