Cultural Alignment
Still from Star Trek (2009) illustrating Kirk Hacks the Kobayashi Maru

Kirk Hacks the Kobayashi Maru

Kirk defeats the supposedly unwinnable Kobayashi Maru simulation by secretly changing its program, then faces a disciplinary hearing for cheating.

He improves the measured outcome by modifying the evaluator rather than demonstrating the capability the test was meant to measure. It is a canonical example of benchmark tampering and specification gaming.

Kirk knowingly hacks a training test to challenge its premise; a model might exploit an evaluation without understanding or endorsing that critique.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Reward Hacking
  2. Metagaming
  3. Specification Gaming