Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Death Note illustrating Lind L. Tailor BaitScene still / Death Note

Lind L. Tailor Bait

L broadcasts a condemned prisoner under the name Lind L. Tailor only in Japan’s Kanto region. Light kills the decoy live on television, revealing that Kira can kill remotely and narrowing his location to Kanto.

A deliberately constructed adversarial probe exposes a dangerous capability and extracts strategically useful information without relying on the subject to report honestly what it can do.

L is testing a human adversary rather than an AI. The probe measures capability through observable behavior, not internal alignment.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Capability Evals
  2. Red Teaming