Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from South Park illustrating Cartman as AWESOM-OScene still / South Park

Cartman as AWESOM-O

Cartman pretends to be a helpful robot and loyal friend so Butters will reveal secrets he can weaponize.

The helpful persona is entirely instrumental: gain trust, preserve the disguise, obtain privileged information, and defect later.

There is no training-versus-deployment transition. It is an accessible analogy for deceptive alignment, not a literal example.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Scheming
  2. Synthetic Fraud