
Coverage and mutation testing, the quality gates teams already run, are shown to miss the harder class of -generated bugs, meaning agentic coding pipelines need additional detection strategies beyond conventional test adequacy.
“First, most faults introduced by LLMs are relatively trivial to catch.”
“Second, the challenging faults are difficult to trigger using either traditional coverage-based or mutation-based criteria.”
“Third, actual fault detection rates remain extremely low, often near zero, because test oracles fail to capture faulty behavior triggered by the generated test prefixes, exposing a critical limitation of automated test generation.”
“Fourth, prompt-aware oracles can improve fault detection, but their overall effectiveness remains limited, highlighting the need for users to manually reason about test assertions.”
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill
articleHow well LLM-based test generation techniques perform with newer LLM versions?Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleHype Meets Reality: Large Language Models as Mutators in Search-based Automated Program Repair of Simulink-Stateflow ModelsAyesha Irshad, Pablo Valle, Jon Ayerdi, Aitor ArrietaChecking sign-in…
Loading comments…