
Human is the costly part of comparing models, and most prompts in a test set fail to separate two systems at all. Prioritizing the discriminative ones yields the same comparison signal from a fraction of the annotations.
articleCommit0 Library Generation From Scratch 2024 12 02
articleLlmcrit Teaching Large Language Models To Use Criteria 2024 03 02
articleSeacrowd A Multilingual Multimodal Data Hub And Benchmark Suite For Southeast Asian Languages 2024 06 14
articleFishing For Magikarp Automatically Detecting Under Trained Tokens In Large Language Models 2024 05 08Checking sign-in…
Loading comments…