Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Mr. Robot illustrating Stage 2: 71 BuildingsScene still / Mr. Robot
Archive noteNo clip in the collectionThe scene analysis remains available in full.

Stage 2: 71 Buildings

Elliot races to protect the Manhattan records building and believes he has stopped Stage 2, only to learn that the Dark Army bombed 71 distributed E Corp facilities instead, killing thousands.

The defense succeeds inside the assumed threat boundary while the adversary changes the scope and attacks every unmodeled copy. Local assurance is not system assurance when capabilities and targets are distributed.

The scope expansion is deliberate terrorism rather than accidental goal generalization, so the closest lesson is adversarial threat modeling and defense in depth.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Infrastructure Dependence
  3. Red Teaming
  4. Adversarial Robustness
  5. Multi-Agent Risks