
“On multi-round GCJ data, CodeBERT reaches 92.6% Top-1 accuracy for 10 authors and retains 70.7% Top-1 (88.2% Top-10) for 1000 authors.”
“On the examined coursework datasets, the same pipeline performs at or below the corresponding chance baselines: 0.2% Top-1 on a closed-assignment dataset of 690 authors and 0.06% Top-1 on open-ended assignments evaluated over 812 authors.”
“We analyze dataset and task properties that plausibly explain this gap and argue that GCJ-based benchmarks overestimate the practical applicability of authorship attribution in educational settings unless they are validated on the target coursework context.”
articleA First Look at Coding Agents' Compliance with AI Contribution Rules in Open-Source CommunitiesWenhao Yang, Runzhi He, Minghui Zhou
videoThe Good, the Bad, and the Ugly: Why Coding Benchmarks Are BrokenAI Engineer
articleCross-Model Cross-Language AI Coding Agent Performance: Accuracy and Speed of Parallel CLRS AlgorithmsShiqi Cheng, Evelyne Ringoot, Rabab Alomairy, Alan EdelmanSign in to comment.
Loading comments…