
Demonstrates that iterative -judge feedback can close much of the quality gap between cheap and expensive drafting agents on a complex professional task, and validates the judge's scores against a human expert, a useful pattern for building agentic document-drafting .
articleVibeCheck: Assessing the Quality of LLM-Generated Unit Tests: A Multi-agent Empirical Study across Heterogeneous RepositoriesAnika Tabassum, Mushahid Intesum, Md. Fahim Arefin, Tarannum Shaila Zaman
articleXAI-Arena: Can LLMs Assess the Quality of XAI Explanations?Yanfei Hu Fleischhauer, Alona Zharova, Nadja Klein, Stefan Feuerriegel
articleGAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented AgentsUmesh Bodhwani, Thanh Tran, Kai WeiChecking sign-in…
Loading comments…