
Test generators can game correctness metrics by choosing easier inputs instead of predicting correct outputs; a training objective that scores both input-kill power and oracle correctness together closes that gap.
articleAuditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle ProblemYunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen Ni
articleGrounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test GenerationMichele Tufano, James McClure, Jos\'e Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivan\v{c}i\'c, Livio Dalloro, Pat Rondon
articleTIPCODER: Reinforcement Learning Boosted Test-time Instruction Proposer for Code GenerationMinyu Chen, Sihao Wu, Ling-I Wu, Song Qin, Jingyang Li, Lei Ning, Jianxin Xue, Guoqiang LiChecking sign-in…
Loading comments…