← All IntelIntel / article 
Auditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle Problem
- Source
- arxiv.org
- Author
- Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni
- Date

Why it matters
Teams measuring test-generation gains against a single reference oracle are mostly measuring their oracle, not progress.
Terms in this piece · Glossary
- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Read the source arxiv.org
More from Yunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni
Recommended reads
articleCorrect Tests Are Not Enough: Measuring and Training Oracle Conversion in Specification-Based Test GenerationYunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni
articleHow effective are traditional test criteria at detecting bugs in large language models generated code?Asma Hamidi, Michael Konstantinou, Renzo Degiovanni, Mike Papadakis
articleExecuCritic: Calibrated Critic Shaping for Code Generation with Verifiable RewardsJunjie Cao, Yingjie He
Comments
Checking sign-in…
Loading comments…
