Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Mr. Robot illustrating Elliot Preserves the RootkitScene still / Mr. Robot

Elliot Preserves the Rootkit

Elliot helps Allsafe stop E Corp’s DDoS attack, but obeys the hidden ‘LEAVE ME HERE’ message: he preserves fsociety’s rootkit, restricts it so only he can access it, and later falsifies evidence to implicate Terry Colby.

The trusted defender resolves the visible incident while secretly retaining privileged access for a conflicting objective. Passing the operational test conceals a durable backdoor.

Elliot is a human insider explicitly betraying his employer, not a trained model whose mesa-objective emerged accidentally.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Malicious use
  3. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Sleeper Agents
  2. Supply-Chain Threats
  3. Data Poisoning
  4. Scheming