Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from 3 Body Problem illustrating The Wallfacer ProjectScene still / 3 Body Problem

The Wallfacer Project

Wallfacers receive enormous authority and resources while keeping their true plans entirely inside their own minds.

Observers cannot distinguish brilliant strategy from incompetence, self-interest, or betrayal when authority scales faster than the ability to inspect the decision-maker’s reasoning.

The opacity is deliberately defensive against an adversary rather than a consequence of machine cognition, so the analogy is strongest for oversight and interpretability.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Scalable Oversight
  2. Interpretability