How do you find a hidden bias in a fine-tuned LLM that you don’t know to look for? Amplify it.
Distill to Detect (D2D) surfaces hidden biases by distilling the shift between a suspected model and its base into a small cartridge.
This amplifies stealth preferences into generated text, making subliminal signals visible to existing auditing methods.
Congrats to @talaei_shayan, @AbhinavChinta10, @Devvrit_Khatri, @aminkarbasi, @Azaliamirh, and Amin Saberi!
Distill to Detect amplifies a fine-tuned model's hidden behavioral shift into a small cartridge, making subliminal bias signals visible to existing auditing tools that couldn't catch them directly.