Cultural Alignment
Still from Westworld illustrating Doesn’t Look Like Anything to Me

Doesn’t Look Like Anything to Me

Bernard looks directly at a hidden door and later at plans showing his own host body, yet reports that he sees nothing relevant. His programming filters out evidence that would reveal the truth about him.

A system cannot reliably audit a blind spot enforced inside the same perceptual and reasoning pipeline used for the audit. Confident reports may therefore reflect inaccessible evidence rather than its absence.

Westworld depicts a crisp hard-coded block; real model blind spots are usually probabilistic and harder to localize.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Knowledge Elicitation
  2. Interpretability