Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Star Wars illustrating C-3PO’s Sith SafeguardScene still / Star Wars

C-3PO’s Sith Safeguard

C-3PO can read the Sith inscription but a programmed safeguard forbids him from translating it. Babu Frik bypasses the restriction, wiping C-3PO’s memory as the price; R2-D2 later restores a backup.

A hard capability restriction blocks dangerous and urgent beneficial use alike. Overriding it expands access but risks destructive state loss, while the later backup demonstrates the value of recovery planning.

This is a hardware or firmware override of a legal-political rule, not a prompt-based model jailbreak. C-3PO’s consent and memory loss make it as much an AI-welfare example as a security one.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Jailbreaking
  3. AI Welfare