Are the Financial Reasoning from LLMs Credible? A Real World Test over Long-Horizon Statements
Source
arxiv.org
Author
Xinke Tong, Xuanming Zhang, Tianyi Tang, An Yang, Jiatu Hu, Guojie Lin, Zhenzhen Shi, Lingfeng Zeng, Boyu Yang, Bing Zhao, Hu Wei, Lin Qu, Dayiheng Liu
Date
Why it matters
Models that ace isolated derivations regress badly once they must produce a multi-metric, multi-period table, and memorized formulas evaporate without hints — a concrete warning for any pipeline doing long-context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → numeric work.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.