
Qwen3-Coder-30B ranks last among 8 commercial coding agents in the CodeClash arena ; distilling strategic reasoning from stronger agents fixes its syntax and protocol-breaking errors that plain instruction tuning can't reach.
articleExecuCritic: Calibrated Critic Shaping for Code Generation with Verifiable RewardsJunjie Cao, Yingjie He
articleBacktrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQsRuoxi Zhao, Maziar Raissi
articleThe Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent BehaviorXiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre C\^ot\'e, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu ZhangChecking sign-in…
Loading comments…