
Tests whether agents can productively prepare a brand-new environment without any task examples, finding a meta- approach beats fixed preprocessing heuristics on 5 of 6 benchmarks, relevant to anyone bootstrapping agents into new codebases or tools.
articleSubagents vs Agent Skills: Executing Reusable Knowledge for Long-Horizon Agentic TasksWasu Top Piriyakulkij, Rachel Lawrence, Alicia Curth, Sushrut Karmalkar, Niranjani Prasad
articleScientific Agent Skills: A Library of Procedural Knowledge for Research AgentsTimothy Kassis, Vinayak Agarwal, Yuhuan He, Darshil Patel, Aubrey M. Brueckner
articleBacktrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQsRuoxi Zhao, Maziar RaissiChecking sign-in…
Loading comments…