Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Star Wars illustrating Galen Erso’s Death Star FlawScene still / Star Wars

Galen Erso’s Death Star Flaw

Forced to finish the Death Star, Galen uses privileged design access to place a hidden weakness in its reactor system, then sends the Rebels a message explaining how the stolen plans can reveal and exploit it.

A trusted insider embeds a catastrophic vulnerability in a high-assurance system, and the organization’s secrecy and weak independent review allow it to reach deployment.

The sabotage is morally beneficial and belongs to conventional engineering, not learned-model behavior. It is an exploitable design vulnerability rather than a trigger-conditioned sleeper policy.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Supply-Chain Threats