Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from True Detective illustrating Rust Spots the EvaluationScene still / True Detective
Archive noteNo clip in the collectionThe scene analysis remains available in full.

Rust Spots the Evaluation

During his recorded 2012 interview, Rust asks whether the conversation is supposed to be admissible, insists on beer, and treats the detectives’ questions as evidence about what they suspect. The apparent witness begins evaluating his evaluators.

A subject who recognizes the test can optimize the impression it gives while using the test itself to infer the evaluator’s hidden information. That is the core danger behind evaluation-aware systems.

Rust is a sophisticated human suspect, and the common claim that drinking automatically makes the interview inadmissible is legally simplistic; the stronger analogy is his strategic shift once he recognizes the evaluation context.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Knowledge Elicitation
  2. Metagaming
  3. Scheming
  4. Situational Awareness