
Gives a measurable criterion for choosing between self-consistency voting and hidden-state answer selection on a given task, and warns that probe accuracy reported without leakage controls is inflated.
articleWhen LLM judges agree, should we believe them?www.amazon.science
articleMemorization Diagnostics for Code LLMs Should be Scale-AwarePrateek Kumar Rajput, Abdoul Aziz Bonkoungou, Alberick Euraste Djir\'e, Xunzhu Tang, Yewei Song, Iyiola Emmanuel Olatunji, El Hacen Diallo, Jacques Klein, Tegawend\'e F. Bissyand\'e
articleMeasure, Don't Optimize: Forecasting Recovery in LLM UnlearningZirui Song, Huaxing Liu, Xiang Wang, Shuai Li, Xinye Li, Lang Gao, Jinghui Zhang, Zheng Lu, Fengxian Ji, Xiaojun Chang, Xiuying ChenChecking sign-in…
Loading comments…