
If a model can tell it is being evaluated, your numbers may not describe deployment. This gives a plug-in way to measure both model awareness and how detectable your own is.
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleOptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentSeonglae Cho, Adriano Koshiyama
postAi2 uses psychometrics to audit what LLM safety benchmarks measureAi2Checking sign-in…
Loading comments…