
Test quality is the usual objection to letting an write tests, and this puts a measured, individually-scored answer against two mature human corpora.
articleAI Evaluation Should Work With HumansJan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleDesigning AI-resistant technical evaluationsAnthropic EngineeringSign in to comment.
Loading comments…