Vibeleaderboard
← All Intel
Intel / article

How Do Coding Agents Optimize Software and Report Performance Validation? A Large-Scale Empirical Study of Open-Source Pull Requests

Source
arxiv.org
Author
Huiyun Peng, Ricardo Calvo, Kelechi G. Kalu, James C. Davis
Date
Why it matters

Agent-authored performance PRs get merged about 19 points less often than human ones despite similar optimization and validation reporting. Useful for anyone letting coding agents open optimization PRs.

Key takeaways · AI-distilled
  • -authored performance PRs leaned more on static reasoning than measurement: 46.1% used benchmarks versus 57.2% of human PRs, though the authors say use converged by the end of the study period.
  • Agents reported validation for code-smell refactorings as often as for resource-targeting changes, while human authors validated refactorings less often than changes aimed at resources.
  • Across both agent and human PRs, about half of the PRs that reported validation included no quantitative performance metric, and trade-offs across performance objectives were rarely quantified.
  • The authors recommend review criteria that check whether a PR measures its intended effect and its costs, rather than treating the mere presence of validation as sufficient. The sample is 1,130 PRs per group, collected through June 2026.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • calibration — How well a model's confidence matches reality — a calibrated model saying "90% sure" is right about 90% of the time.
  • benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Recommended reads
Comments

Checking sign-in…

Loading comments…