
Quantifies how often off-the-shelf LLMs falsely flag or miss requirements defects versus expert ground truth, showing teams should validate rather than trust default requirements review.
articleDo large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effectsEmad Alharbi
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleBeyond the Traceback: Using LLMs for Adaptive Explanations of Programming ErrorsAlexandru-Radu Moraru, Shreyan Biswas, Ujwal GadirajuChecking sign-in…
Loading comments…