DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?
Source
Susana Haing, Natan Vidra, Spurthi Setty
Author
Susana Haing, Natan Vidra, Spurthi Setty
Date
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
It quantifies something every team using coding agents feels but does not measure: green tests are a weak proxy for the code doing what you actually meant.