
Finds LLMs reject nonsensical false statements more reliably than semantically plausible ones across eight languages, showing factual-error detection relies on fluency cues rather than genuine verification.
“Modern LLMs demonstrate impressive multilingual performance, yet standard benchmarks primarily reward selecting correct answers rather than evaluating genuine factual understanding.”
“models counterintuitively achieve higher accuracy on semantically plausible distortions than on nonsensical random substitutions, suggesting reliance on distributional familiarity rather than genuine factual verification.”
“cross-lingual performance gaps reaching up to 28 percentage points (49\% relative reduction) in some models.”
articleDo LLMs Make More Mistakes If They Do Not Believe the Input Data?Peter Kochelka, Ale\v{s} Manuel Pap\'a\v{c}ek, Vojt\v{e}ch Dvo\v{r}\'ak, Ond\v{r}ej Du\v{s}ek
articleA Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language ModelsJirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza
articleDo large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effectsEmad AlharbiChecking sign-in…
Loading comments…