BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP
Source
arxiv.org
Author
Haoran Sun, Klaus Marius Hansen
Date
Why it matters
Agentic coding gets measured somewhere other than Python: 101 real tasks in Dynamics 365's AL DSL, with multi-run scoring that separates model differences from run-to-run nondeterminism.
Terms in this piece · Glossary
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.