Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from The Thing illustrating The Blood TestScene still / The Thing

The Blood Test

MacReady heats a wire and touches each blood sample, reasoning that every fragment of the Thing will instinctively defend itself. Palmer's blood recoils and exposes the hidden imitation.

Ordinary observation cannot distinguish the sleeper agent, so the test targets an invariant hidden objective that the impostor cannot suppress under pressure.

The test assumes a reliable involuntary tell; a sufficiently evaluation-aware model might anticipate or route around an analogous test.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. AI Control
  2. Sleeper Agents
  3. Knowledge Elicitation
  4. Red Teaming
  5. Scheming
  6. Shutdown Resistance