Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Hunter × Hunter illustrating Kurapika’s Chain JailScene still / Hunter × Hunter
Archive noteNo clip in the collectionThe scene analysis remains available in full.

Kurapika’s Chain Jail

Kurapika makes Chain Jail overwhelmingly powerful by restricting it to members of the Phantom Troupe; using it on anyone else will kill him.

A narrowly scoped capability becomes safer and stronger because the restriction is technically binding and self-enforcing, not merely a policy the user can ignore when tempted.

The penalty is magical and perfectly enforceable. Real AI controls are imperfect, gameable, and vulnerable to distribution shift.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. AI Control
  2. Safe Exploration
  3. Behavioral Alignment