
Scaffolding tuned to older models can become dead weight — or a regression — when the model underneath changes.
“In other words, stronger (newer) LLMs may obviate any advantage these techniques bring.”
“Our results show that the plain LLM approach can outperform previous state-of-the-art approaches in all test effectiveness metrics we used: line coverage (by 17.72%), branch coverage (by 19.80%) and mutation score (by 20.92%), and it does so at a comparable cost (LLM queries).”
“This strategy achieves comparable (slightly higher) effectiveness while requiring about 20% fewer LLM requests.”
articleHow effective are traditional test criteria at detecting bugs in large language models generated code?Asma Hamidi, Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleTangent: An Empirical Study of Testing Practices for LLM-Based Agent ApplicationsRangeet Pan, Tyler Stennett, Divya Sankar, Bridget McGinn, Alessandro Orso, Raju Pavuluri, Saurabh Sinha, Maja Vukovic
articleScaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model ParametersCharlie Snell et al.Checking sign-in…
Loading comments…