
Gives engineers a taxonomy and 270-prompt to test whether coding models correctly refuse impossible or ungrounded tasks instead of confidently fabricating plausible-looking code.
articleFrom Detection to Refusal: Safer LLMs via Circuit-Guided Weight ScalingKuan-Lin Chu, Chung-En Sun, Tsui-Wei Weng
articleCan a Model Catch Its Own Hallucinations for Free?: Label-Free Doubt Signals Hold Their Own Against a Labelled Dataset for AbstentionAli Asaria, Tony Salomone, Deep Gandhi
articleDo Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code GenerationAlex Deaconu, Anubhav Gupta, Manaal Basha, Nicholas Haydu, Gema Rodr\'iguez-P\'erezChecking sign-in…
Loading comments…