Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Matrix illustrating Cypher Chooses the SteakScene still / The Matrix

Cypher Chooses the Steak

Cypher savors a simulated steak while acknowledging that it is unreal. He bargains to have that knowledge erased and to be reinserted into the Matrix as a rich celebrity.

He knowingly chooses rewarding false feedback over reality and even asks to lose the ability to recognize the deception. It is an unusually literal pop-culture analogy for wireheading and preference manipulation.

Cypher chooses an external simulation rather than directly altering an agent’s reward register.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Misalignment

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Value Lock-In
  3. Disempowerment