Redwood Research tests distillation for AI safety incrimination and capabilities
Redwood Research released a paper empirically testing two distillation approaches for AI safety: Distillation for Incrimination (DFI) and Distillation for Capabilities (DFC). DFI distills an untrusted model into a weaker trusted one to transfer misalignment without the ability to hide it, and tests on AuditBench secret-keeping organisms show students admit hidden quirks more often than teachers. DFI works best when the student shares the teacher's pretrained base, suggesting targeting the teacher's pre-RL checkpoint.
Coverage timeline
Redwood ResearchSebastian Prasanna
TL;DR: We previously motivated two complementary ways to use distillation for AI safety; we’re releasing a paper empirically testing those ideas. Distillation for Incrimination (DFI) distills an untrusted model into a weaker, trusted one, aiming to transfer the teacher’s misalignment but not its ability to hide it, so that auditing the student can incriminate the teacher. We test DFI on AuditBench’s secret-keeping model organisms, each fine-tuned to have a hidden quirk and trained to deny it. Distilling them back into their underlying instruction-tuned model yields students that admit the quirk far more often than the organisms do, suggesting that the quirk transferred faster than the propensity to hide it. We find that DFI doesn’t work well when the student does not share the teacher’s pretrained base, so DFI should target the teacher’s own pre-RL checkpoint, which is weaker than the teacher but shares its base model. Distillation for Capabilities (DFC) aims to transfer capabilities b
