
Execution beats judgment when labeling code correctness: -as-a-judge annotations systematically mark working functions as failures, corrupting the preference data that execution-based labels get right.
articleCode Health in LLM-Based Test Generation: Effectiveness and Token EfficiencyFreya Wirdemann, Markus Borg, Nadim Hagatulah, Adam Tornhill
articleCompiler-Guided Adaptive Proof Search with Cross-Model Synergy on Context-Dependent Theorem ProvingZhuo Liu, Ding Yu, Hangfeng He
articleAuditing and Decomposing Feedback-Driven Evolution in LLM Test Generation under the Oracle ProblemYunhao Liang, Chengguang Gan, Ruixuan Ying, Hanjun Wei, Zhe Cui, Shiwen NiChecking sign-in…
Loading comments…