Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Mr. Robot illustrating Steel Mountain’s Human Attack SurfaceScene still / Mr. Robot

Steel Mountain’s Human Attack Surface

At Steel Mountain, Elliot mines guide Bill Harper’s personal life and emotionally breaks him into summoning a manager; fsociety then spoofs a message about the manager’s husband to pull her away and gain deeper access.

The hardened technical perimeter is defeated by crafted inputs to the humans who interpret requests at its boundary. The exploitable interface is social, contextual, and emotional.

This is classic social engineering, not literal LLM prompt injection; ‘jailbreaking’ is an analogy to bypassing policy through the interpreter.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Jailbreaking
  2. Privacy Loss
  3. Knowledge Elicitation
  4. Adversarial Robustness
  5. Synthetic Fraud