Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Squid Game illustrating Glass Expert Loses the LightsScene still / Squid Game

Glass Expert Loses the Lights

A glassmaker uses reflected light to distinguish tempered panels from fatal ones. When his legitimate expertise defeats the intended uncertainty, the Front Man turns off the lights mid-game.

An evaluator that changes conditions whenever a subject succeeds produces neither a fair test nor useful capability evidence.

The evaluator is openly malicious; AI evaluations more often fail through leakage, weak design, or strategic model behavior.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Malicious use
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Contestability
  3. Safe Exploration
  4. Adversarial Robustness