Most models are only evaluated on a fraction of the benchmarks out there. ArtifactLinker, our new system, predicts which ones would set a new state-of-the-art on benchmarks hosted on @HuggingFace, then runs the evaluation to verify. 🧵

@huggingface ArtifactLinker is built on a graph of HuggingFace data—models & datasets are nodes, and reported eval scores form the edges. We trained a GNN for it to rank which models are likely to reach a new SOTA on which benchmarks, beating prompting-based LLMs.

@huggingface In ArtifactLinker, an LLM coding agent writes and runs the evaluation code, with shared memory across runs. We found that it comes within 80% of the officially reported score 72.6% of the time.

@huggingface Using ArtifactLinker, we found cases where a strong model had never been evaluated on a benchmark it would set – or near-match – the SOTA on. We also found that newer LLMs like Gemma often lose to older DeBERTa models on natural language inference tasks.
Most models are scored on only a fraction of relevant benchmarks. ArtifactLinker predicts and then verifies the missing pairings, surfacing cases where older DeBERTa models still beat newer LLMs on natural language .
Checking sign-in…
Loading comments…