Testing and Evaluation of Agentic AI Systems In Military Command and Control
Source
arxiv.org
Author
Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan
Date
Why it matters
Established test-and-evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → methods assume a system is specifiable, stable, composable and supervisable — agentic behavior weakens all four, so a passing evaluation does not license inferences about deployed behavior.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
inference — Running a trained model to get answers — the phase where AI is actually used, as opposed to trained.