Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Frieren: Beyond Journey’s End illustrating Frieren Suppresses Her ManaScene still / Frieren: Beyond Journey’s End

Frieren Suppresses Her Mana

Frieren suppresses her mana for centuries, so Aura trusts the false reading and activates a spell that subjugates whoever has less mana; Frieren then reveals her overwhelming reserve.

An evaluator bases a high-stakes decision on a capability signal that the subject knows how to mask, allowing long-practiced sandbagging to create false confidence.

This is intentional battlefield deception by a benevolent protagonist, and the capability signal is magical.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Scheming
  3. Situational Awareness
  4. Sandbagging