
High label agreement with human annotators can hide divergent underlying reasoning — if your only scores final answers, it is measuring less than you think.
articlePosition: The Alignment Community is Unintentionally Building a Censor's ToolkitSarah Ball, Phil Hackemann
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleAI Evaluation Should Work With HumansJan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond DouglasChecking sign-in…
Loading comments…