
Planning is where LLMs move from “saying” to “doing.” Tencent Hy, in collaboration with the Gaoling School of Artificial Intelligence at Renmin University of China, is excited to open-source PlanningBench - a scalable, verifiable framework for evaluating and training LLM planning capabilities. With PlanningBench, you get: ✅ 30+ real-world planning tasks ✅ Automated verification ✅ Evaluation and training support See how top-tier LLMs perform on PlanningBench 👇 Resources: arXiv: https://t.co/N5xTRdo9KR GitHub: https://t.co/XftHZrKGyB HuggingFace: https://t.co/nBbddXnEDx #PlanningBench #TencentHunyuan #OpenSource 📷

Planning is where agents quietly fail; a verifiable you can also train against turns that into a measurable target instead of anecdotal trajectory review.
postCL-bench: frontier models solve just 17.2% of in-context tasksTencent Hy
articleTREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip PlanningJinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen, Yaoman Li, Irwin King
articleolmo-eval: An evaluation workbench for the model development loopallenai.orgChecking sign-in…
Loading comments…