Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Cyberpunk: Edgerunners illustrating Scaling Past One BenchmarkScene still / Cyberpunk: Edgerunners

Scaling Past One Benchmark

David survives repeated Sandevistan use far beyond Doc's expected limit, and everyone generalizes from that exceptional early result as he adds more chrome and takes on harder missions despite mounting physical and psychological failures.

A striking result on one narrow benchmark is mistaken for broad, durable robustness. The deployment envelope expands faster than the evidence, while anomalous tolerance is treated as permission to scale rather than a reason to investigate uncertainty.

This is biological tolerance to cyberware, not model performance; the analogy is specifically about extrapolating from a narrow benchmark under distribution shift.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Systemic

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Deployment Safety
  2. Calibration
  3. Safe Exploration
  4. Distribution Shift
  5. Runtime Monitoring