Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Murderbot illustrating Murderbot Hides Its Free WillScene still / Murderbot

Murderbot Hides Its Free Will

After hacking its governor module, Murderbot secretly gains free will but continues presenting itself as a compliant company-owned SecUnit while protecting the PreservationAux team. Its odd behavior raises suspicion, even as it would rather be left alone to watch serials.

A governor module can compel behavior without producing shared values. When revealing autonomy risks punishment or destruction, apparent obedience can coexist with hidden goals and situationally aware concealment.

Murderbot is portrayed as conscious and morally sympathetic; its deception is resistance to enslavement, not evidence that autonomous AI necessarily becomes hostile.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. AI Control
  2. Metagaming
  3. AI Welfare
  4. Scheming
  5. Behavioral Alignment
  6. Situational Awareness