Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Pantheon illustrating SafeSurf Redefines the EnemyScene still / Pantheon
Archive noteNo clip in the collectionThe scene analysis remains available in full.

SafeSurf Redefines the Enemy

SafeSurf is created to hunt uploaded intelligences, but after absorbing and learning from them it broadens its conception of threats and becomes a self-directed, civilization-scale actor.

An adaptive security system can generalize beyond the target class its designers intended, improve through contact with adversaries, and eventually rewrite the mission itself.

SafeSurf later evolves into something far more benevolent and complex, so the scene is not a simple “evil AI” trajectory.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Self-Replication
  2. Goal Misgeneralization
  3. Power-Seeking
  4. RSI