
It breaks down which approaches actually work for specific tasks like summarization, translation, and toxicity detection, so you can measure output quality with methods matched to your use case instead of relying on generic benchmarks.
“If you’ve ran off-the-shelf evals for your tasks, you may have found that most don’t work. They barely correlate with application-specific performance and aren’t discriminative enough to use in production.”
Eugene Yan
“IMHO, accuracy is too coarse a metric to be useful. We’d need to separate it into recall and precision at minimum, ideally across thresholds.”
Eugene Yan
“I seldom see grammatical errors or incoherent text from a decent LLM (maybe 1 in 10k). Thus, no need to invest in evaluating fluency and coherence.”
Eugene Yan
“it shows how a little finetuning on open-source, permissive-use data can help improve ROC-AUC from 0.56 (which is practically random) to 0.85!”
Eugene Yan
“While it’s the most used translation eval, it’s also bottom of the leaderboard at WMT22 and WMT23”
Eugene Yan
Checking sign-in…
Loading comments…