Measuring the Cross-Lingual Comprehension Gap: How the language of the evidence shapes what language models understand
Source
arxiv.org
Author
Rafael da Silva, Jeff Eicher
Date
Why it matters
Holding content, question and model fixed, answer quality drops about 17% relative when the passage is in a non-English language, and worse for lower-resource ones — a concrete tax to budget for before shipping multilingual retrieval or support features.
Terms in this piece · Glossary
token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.