
Human oversight of agents degrades under exactly the conditions agents create — the paper names design affordances and org protocols that keep overseers exercising judgment instead of rubber-stamping.
articleAI Evaluation Should Work With HumansJan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas
articlePosition: Behavioral Systems Require Behavioral TestsManuel Cherep, Nikhil Singh, Pattie Maes
articlePosition: Multi-Agent Systems Should Prioritize Concurrency ControlXin Yang, Letian Li, Zimo Ji, Terry Jingchen Zhang, Wenyuan JiangChecking sign-in…
Loading comments…