How to evaluate long- Q&A systems: datasets, methodology and the pitfalls that make scores lie.
“Positional variance: Evidence may appear at the beginning, middle, or end of documents, making it a challenge for models with limited effective context or those susceptible to the “lost in the middle” problem.”
Eugene Yan
“We also want to distinguish faithfulness from correctness. An answer might be correct based on general knowledge but still be unfaithful if it contradicts the document.”
Eugene Yan
“A study by Xu et al. (2023) found that domain experts in fields like biology or economics preferred answers that were both comprehensive and faithful, particularly for long-form questions. In contrast, crowd-workers often emphasized surface aspects such as conciseness or detail.”
Eugene Yan
“NarrativeQA intentionally generates questions based on summaries rather than full texts. This encourages questions that test narrative comprehension rather than shallow fact recall. For the same reason, QASPER creates questions based on abstracts from academic papers that models then answer based on the full paper.”
Eugene Yan
“People find comparing two answers easier than assigning absolute ratings, resulting in greater consistency across annotators.”
Eugene Yan
articlePrompting Long ContextAnthropic News
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
postAi2 uses psychometrics to audit what LLM safety benchmarks measureAi2Checking sign-in…
Loading comments…