
Standard recall metrics overstate protection: prompts safety monitors miss are up to 5.6x more likely to get a compliant response from the target model than the ones they catch.
“Safety monitors screen prompts sent to deployed language models, flagging harmful requests so they are never answered.”
“The prompts a monitor misses are 2.8 to 5.6 times more likely to be complied with than the prompts it catches.”
“This suggests that standard recall may overstate the protection monitors provide in practice, and that monitors should be evaluated against what their models will actually answer.”
articleFrom Detection to Refusal: Safer LLMs via Circuit-Guided Weight ScalingKuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
articleDon't Want Your LLM to Recommend Nuclear Strike? Try Asking It in JapaneseRian Touchent (ALMAnaCH)
articleMemorization Diagnostics for Code LLMs Should be Scale-AwarePrateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'eChecking sign-in…
Loading comments…