
Standard recall metrics overstate protection: prompts safety monitors miss are up to 5.6x more likely to get a compliant response from the target model than the ones they catch.
articleFrom Detection to Refusal: Safer LLMs via Circuit-Guided Weight ScalingKuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
articleDon't Want Your LLM to Recommend Nuclear Strike? Try Asking It in JapaneseRian Touchent (ALMAnaCH)
articleMemorization Diagnostics for Code LLMs Should be Scale-AwarePrateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'eChecking sign-in…
Loading comments…