
If you use LLMs as judges or annotators, their scores carry measurable severity and biases. This gives a psychometric way to diagnose those biases instead of trusting raw agreement with humans.
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleInter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended TextHaoyuan Li, Snigdha Chaturvedi
articleThe Limits of Automatic Evaluation of Creativity in Large Language ModelsAlessandro Tutone, Giorgio Franceschelli, Mirco MusolesiChecking sign-in…
Loading comments…