Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Cowboy Bebop illustrating Chessmaster Hex’s Sleeper PayloadScene still / Cowboy Bebop

Chessmaster Hex’s Sleeper Payload

A fired gate-system developer plants sabotage intended to fire after the network’s next major upgrade. Corporate stagnation delays activation for decades, until Hex is senile and has forgotten making it.

A dormant backdoor can outlive its author’s intent, memory, and original political context while remaining embedded in critical infrastructure.

This is deterministic human-written sabotage, not a learned model strategically concealing an objective.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Sleeper Agents
  3. Supply-Chain Threats
  4. Long-Horizon Autonomy