
Refusal is not one monolithic behavior. Knowing which heads and neurons carry the safety signal gives a concrete target for hardening a model against jailbreaks without retraining or changing its architecture.
articleDo large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effectsEmad Alharbi
articleMemorization Diagnostics for Code LLMs Should be Scale-AwarePrateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'e
articleHype Meets Reality: Large Language Models as Mutators in Search-based Automated Program Repair of Simulink-Stateflow ModelsAyesha Irshad, Pablo Valle, Jon Ayerdi, Aitor ArrietaChecking sign-in…
Loading comments…