Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Fullmetal Alchemist: Brotherhood illustrating Shou Tucker’s Final AssessmentScene still / Fullmetal Alchemist: Brotherhood

Shou Tucker’s Final Assessment

Facing loss of his certification and livelihood, Tucker creates another “successful” talking chimera by combining Nina and Alexander.

A capability assessment becomes the target, rewarding an impressive output while ignoring provenance, consent, and catastrophic externalities.

Tucker knowingly commits an atrocity; this is chiefly an institutional-incentive failure rather than accidental machine misalignment.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Goodhart’s Law
  3. Governance Failure