Function-Level Execution Feedback for Code Preference Optimization
Source
arxiv.org
Author
Idris Nechnech, Sehwan Kim, Jimin Seo, Yeongoon Kim, Minhae Oh, Sangwoo Hong, Jungwoo Lee
Date
Why it matters
Execution beats judgment when labeling code correctness: LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition →-as-a-judge annotations systematically mark working functions as failures, corrupting the preference data that execution-based labels get right.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.