Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Big Short illustrating The Ratings-Agency MeetingScene still / The Big Short

The Ratings-Agency Meeting

Ratings-agency analysts explain that if they refuse to award permissive ratings, issuers will take their business to a more accommodating competitor.

Issuer-paid evaluators compete to provide permissive ratings, which the whole market then treats as authoritative and propagates.

The evaluators are human institutions responding to incentives; this is institutional capture, not a technical compromise of an AI evaluator.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Governance Failure
  2. Multi-Agent Risks
  3. Race Dynamics