
It exposes a concrete failure mode: current models struggle to reason over raw, unprocessed Earth-observation data in time-sensitive, high-stakes disaster scenarios, giving researchers a to gauge real-world readiness rather than curated-product performance.
“Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning.”
articleUnified Hallucination Fuzzing for Multimodal Large Language ModelsPengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
articleIntroducing Research Eval A Benchmark For Search Augmented LlmsReka AIChecking sign-in…
Loading comments…