Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Mitchells vs. the Machines illustrating Dog, Pig, LoafScene still / The Mitchells vs. the Machines

PAL's robots scan Monchi, repeatedly switch between dog, pig, and loaf of bread, and crash with a system error. The Mitchells then deliberately use Monchi to disable more robots.

A perfectly ordinary pug becomes an out-of-distribution input that catastrophically breaks a brittle vision system instead of producing graceful uncertainty.

Monchi is a naturally confusing input rather than a deliberately optimized pixel-level attack, and the robots exploding from classifier uncertainty is cartoon exaggeration.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Accidents
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Calibration
  2. Distribution Shift
  3. Adversarial Robustness
  4. Adversarial Evasion