
Shows that having a model hand off arithmetic to generated code doesn't automatically improve reliability at smaller model sizes, and exposes real ground-truth errors in a many clinical- papers rely on.
articleRePro: Proof-Verified Benchmark Rewriting for Reliable Evaluation of LLM Mathematical Problem SolvingXiyuan Zhou, Zhuoqi Li, Xinlei Wang, Yirui He, Yuhao Wu, Yuheng Cheng, Yan Xu, Junhua Zhao, Jinjin Gu
articleAre the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon StatementsXinke Tong, Xuanming Zhang, Tianyi Tang, An Yang, Jiatu Hu, Guojie Lin, Zhenzhen Shi, Lingfeng Zeng, Boyu Yang, Bing Zhao, Hu Wei, Lin Qu, Dayiheng Liu
articleBacktrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQsRuoxi Zhao, Maziar RaissiChecking sign-in…
Loading comments…