Vibeleaderboard
← All Intel
Intel / article

BC-Bench: Evaluating Agentic Engineering in a Domain-Specific Language for ERP

Source
arxiv.org
Author
Haoran Sun, Klaus Marius Hansen
Date
Why it matters

Agentic coding gets measured somewhere other than Python: 101 real tasks in Dynamics 365's AL DSL, with multi-run scoring that separates model differences from run-to-run nondeterminism.

Terms in this piece · Glossary
  • SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
Recommended reads
Comments

Checking sign-in…

Loading comments…