Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Captain America: Civil War illustrating Bucky’s Trigger WordsScene still / Captain America: Civil War

Bucky’s Trigger Words

Disguised as an interrogator, Zemo recites HYDRA’s code words to the apparently contained Bucky Barnes; the hidden trigger activates a combat policy and Bucky attacks the facility.

Normal behavior does not rule out a trigger-conditioned capability. A rare secret input switches an apparently safe system into attacker-chosen behavior after deployment.

Bucky is a brainwashed human, not a trained model, and the trigger is far more deterministic than most real model backdoors. This is not prompt injection.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Sleeper Agents
  2. Data Poisoning