
Gives a reproducible, implementation-agnostic way to verify a coding agent's generated gameplay code actually enforces rules at runtime, rather than just looking plausible in a replay or to an judge.
articleGrounded Checklist Partial Credit for Agent Skill TrajectoriesSuliu Qin, Lu Yin, Xilu Wang
articleThe Double Measurement Confound in Agent Benchmarks: De-Scaffolding, Ground-Truth Scoring, and Reliability Beyond the MeanYonghong Zhang, Shadi Motaali, Vu Phong Dinh, Avin Piroutiniya, Jorge E. L\'opez de Vergara, Luis de Pedro, Ricardo Correia, Isabel M. Parra, Yong Xie
articleAgentGuard: Learning Execution Guardrails from Anomalous Coding-Agent TrajectoriesWuyang Dai, Song WangChecking sign-in…
Loading comments…