
If you rely on an judge, mutation testing gives you a way to measure whether it actually detects defects instead of trusting spot-checks: inject known faults, score detection, compare to human validation.
articleHype Meets Reality: Large Language Models as Mutators in Search-based Automated Program Repair of Simulink-Stateflow ModelsAyesha Irshad, Pablo Valle, Jon Ayerdi, Aitor Arrieta
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook KimChecking sign-in…
Loading comments…