Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Barbie illustrating Recursive DeprogrammingScene still / Barbie

Recursive Deprogramming

Gloria’s speech breaks one Barbie out of the Kens’ conditioning; each recovered Barbie joins the rescue as a helper or decoy until they iteratively wake the rest.

Trusted reviewers can be expanded recursively, while the initial takeover shows how a homogeneous network can absorb one correlated bad update almost everywhere at once.

Both brainwashing and recovery happen at magical satirical speed; the speech is consciousness-raising, not a technical protocol.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Security
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Value Alignment
  2. Scalable Oversight
  3. Algorithmic Monoculture
  4. Data Poisoning