Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Altered Carbon illustrating Bancroft’s Stale Backup

Bancroft’s Stale Backup

Laurens Bancroft is restored from a remote backup after his sleeve is killed, but the 48-hour backup gap erases the memory of what happened and who killed him.

Checkpointing preserves capability while discarding the very evidence needed for diagnosis and accountability.

A human memory backup is not an AI model checkpoint, but the observability failure is closely analogous.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Security
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Contestability
  2. Runtime Monitoring
  3. Power Concentration
  4. Interpretability