
lm-evaluation-harness v0.4.12
github.com- Category
- Other
- Type
- ARTICLE
- Builder
- EleutherAI
- Added
- Jul 22, 2026
About
New release with four new model backends, tensor parallel support for `transformers` based models (`hf`), new benchmarks, a `TaskManager` refactor, and a long tail of task correctness fixes. ## Highlights ### New Model Backends * **TensorRT-LLM (`trt-llm`)** — NVIDIA TensorRT-LLM backend for optimized GPU inference by @Tracin in #3628 * **Megatron-LM (`megatron-lm`)** — Megatron-LM backend with TP/EP/DP support by @shangxiaokang in #3521 (with follow-up hardening in #3607) * **Intel G
Why it made the leaderboard
If you evaluate open-weight LLMs, this release adds first-class backends for TensorRT-LLM, Megatron-LM, Gaudi, and a LiteLLM gateway plus native multi-GPU tensor parallelism for HF models — letting you benchmark the same tasks across far more inference stacks. Watch the breaking changes (vLLM >=0.18, SteeredHF renamed, stricter enable_thinking rules) before upgrading.
Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.