Vibeleaderboard
Index / article

lm-evaluation-harness v0.4.12

github.com
Visit github.com
Category
Other
Type
ARTICLE
Builder
EleutherAI
Added
Jul 22, 2026

About

New release with four new model backends, tensor parallel support for `transformers` based models (`hf`), new benchmarks, a `TaskManager` refactor, and a long tail of task correctness fixes. ## Highlights ### New Model Backends * **TensorRT-LLM (`trt-llm`)** — NVIDIA TensorRT-LLM backend for optimized GPU inference by @Tracin in #3628 * **Megatron-LM (`megatron-lm`)** — Megatron-LM backend with TP/EP/DP support by @shangxiaokang in #3521 (with follow-up hardening in #3607) * **Intel G

Why it made the leaderboard

If you evaluate open-weight LLMs, this release adds first-class backends for TensorRT-LLM, Megatron-LM, Gaudi, and a LiteLLM gateway plus native multi-GPU tensor parallelism for HF models — letting you benchmark the same tasks across far more inference stacks. Watch the breaking changes (vLLM >=0.18, SteeredHF renamed, stricter enable_thinking rules) before upgrading.

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.