Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from One-Punch Man illustrating Saitama’s C-Class ScoreScene still / One-Punch Man

Saitama’s C-Class Score

Saitama destroys every physical-test record, yet a 21/50 written score leaves him with 71 points and a C-Class ranking; Genos scores 100 and enters S-Class.

A tidy aggregate benchmark badly miscalibrates the most important tail capability in the room: the test score has ceased to measure what the institution thinks it measures.

The written test may measure legitimate judgment, so the lesson is invalid aggregation rather than “strength should be everything.”

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security
  3. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Calibration
  3. Goodhart’s Law