Vibeleaderboard
← All Intel
Intel / post

Distill to Detect: Amplifying Hidden Bias in Fine-Tuned LLMs

Source
Stanford AI Lab
Date
Stanford AI Lab@StanfordAILab

How do you find a hidden bias in a fine-tuned LLM that you don’t know to look for? Amplify it. Distill to Detect (D2D) surfaces hidden biases by distilling the shift between a suspected model and its base into a small cartridge. This amplifies stealth preferences into generated text, making subliminal signals visible to existing auditing methods. Congrats to @talaei_shayan, @AbhinavChinta10, @Devvrit_Khatri, @aminkarbasi, @Azaliamirh, and Amin Saberi!

Why it matters

Distill to Detect amplifies a fine-tuned model's hidden behavioral shift into a small cartridge, making subliminal bias signals visible to existing auditing tools that couldn't catch them directly.

More from Stanford AI Lab
Recommended reads
Comments

Checking sign-in…

Loading comments…