
Reasoning models are increasingly asked to estimate their own chance of success; CALIBER addresses calibrating that estimate both pre- and post-reasoning, which is the signal pipelines use for abstention, retries, and escalation.
articleBam Just Like That Simple And Efficient Parameter Upcycling For Mixture Of Experts 2024 08 15
articleMultilingual Arbitrage Optimizing Data Pools To Accelerate Multilingual Progress 2024 08 28
articleOne Tokenizer To Rule Them All Emergent Language Plasticity Via Multilingual Tokenizers 2025 05 30
articleInclude Evaluating Multilingual Language Understanding With Regional Knowledge 2024 11 29
articleOptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentSeonglae Cho, Adriano Koshiyama
articleConfidence Estimation for Financial Vision-Language Models in Chart and Document UnderstandingReza Khanmohammadi, Simerjot Kaur, Charese H. Smiley, Ivan Brugere, Mohammad M. Ghassemi
articleThe Knowing-Saying Gap: When Probes See Errors that Confidence MissesJyotin Goel, Ipshita Bandyopadhyay, Justin ShenkChecking sign-in…
Loading comments…