Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Toy Story illustrating Buzz’s Lucky FlightScene still / Toy Story

Buzz’s Lucky Flight

Buzz jumps from the banister to prove he can fly. A chain of lucky collisions carries him around the room and back to a perfect landing, convincing the toys—and Buzz—that he passed the test.

An outcome-only evaluation credits capability that actually came from environmental luck. Without inspecting the causal process, one spectacular demo can produce dangerously false confidence.

The sequence is cartoon physics, and Buzz’s false belief is a character gag rather than a trained model’s claim.

AI Risk families

This scenario is an example of this type of AI risk.
  1. Accidents

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Calibration
  3. Distribution Shift
  4. Interpretability