Cultural Alignment

Exploring real AI risks through the lens of pop culture

Still from Harry Potter illustrating Slughorn’s True MemoryScene still / Harry Potter

Slughorn’s True Memory

Slughorn possesses the crucial memory of Tom Riddle asking about multiple Horcruxes, but initially gives Dumbledore a visibly doctored version. Harry changes the elicitation strategy and persuades Slughorn to surrender the complete, faithful memory.

The information exists inside the source, yet a direct request returns a strategically altered answer. Recovering the truth requires a different interaction strategy rather than simply asking the same question more forcefully.

This is human persuasion, not a formal eliciting-latent-knowledge protocol. Slughorn also knows that he is withholding information.

AI Risk families

This scenario is an example of the following types of AI risk.
  1. Misalignment
  2. Security

AI safety concepts

This scenario is related to the following AI safety concepts.
  1. Knowledge Elicitation
  2. Data Poisoning