Vibeleaderboard
Index / tool
Visit github.com
Category
Developer Tools
Rank
Pricing
Open Source
Type
TOOL
Use case
Models: Train & Run
Interfaces
API
Latest release
v0.18.0
Date

About

LMDeploy is an open-source toolkit for compressing, quantizing, and serving large language and vision-language models, offering an OpenAI/Anthropic-compatible API server, a PyTorch and TurboMind inference engine, and weight/KV-cache quantization that claims up to 1.8x higher throughput than vLLM and 2.4x faster 4-bit inference than FP16.

What it does

LMDeploy packages two swappable engines, one hand-tuned in C++ for raw speed and one written entirely in Python for easy hacking, behind a single command-line tool, a Python library, and a hosted API server. Point it at a model on Hugging Face and it picks an engine, shrinks the model's weights and cache to lower-precision numbers, and starts answering requests, batching many users' prompts together instead of running them one at a time. The same install handles ordinary chat models and models that also read images.

Why it's ranked here

OpenMMLab built this to cover ground that CUDA-only servers skip: the same install targets Ascend, Cambricon and MetaX accelerators alongside Nvidia GPUs, docs list Llama, Qwen, InternLM, DeepSeek, GLM, Mixtral and a matching column of vision-language models, and every model routes through one of two engines, a compiled one built for throughput and a pure Python one built so new architectures land without a rebuild. Few open serving projects carry that hardware breadth and model breadth at once, which is why this belongs on a shortlist even before its speed claims are checked independently.

What's good

The two-engine split is well thought out rather than bolted on: the compiled backend reports itself unavailable rather than crashing when its extension isn't built, so the pure Python engine keeps working without it. The command line tool exposes the details that matter for production instead of hiding them behind defaults, quantization policy per request, tensor and expert parallelism for mixture-of-experts checkpoints, and prefix caching. Model coverage is unusually wide for a serving project, spanning ordinary chat models up to trillion-parameter mixture-of-experts systems and their vision-capable counterparts.

Tradeoffs

Only the compiled engine gets full performance, and building it needs a working cmake and pybind toolchain, a compatible CUDA release, and an actual Nvidia GPU; the pure Python engine is the fallback everywhere else, including plain CPU boxes. Even a single-file offline script pulls in the full serving stack, since the top-level import wires the client, pipeline and server together rather than keeping them separate. The extensive model list also only says a model is wired in, not that it is fast: the newest, largest entries assume multi-GPU setups this review cannot verify.

How to use it well

Reach for this when serving a model in production on real GPU hardware, especially Ascend, Cambricon or MetaX chips where fewer serving frameworks even try, or when quantizing a model's weights and cache to fit tighter memory. The command line covers both engines directly, so start there to pick tensor parallelism, expert parallelism for mixture-of-experts checkpoints, and a quantization policy before writing any Python. For quick offline scoring or batch generation without standing up a server, the library's pipeline interface is the shorter path; skip the compiled engine entirely if there's no CUDA toolchain to build it.

Technical notes+

The package entry point is lmdeploy.cli:run, wired from lmdeploy/cli/cli.py's CLI class, whose chat subcommand exposes separate PyTorch-engine and TurboMind-engine argument groups (tensor parallel, expert parallel, cache_max_entry_count, quant_policy, prefix caching, speculative decoding). lmdeploy/turbomind/__init__.py imports the compiled _turbomind extension inside a try/except and exposes is_available(), so the compiled path degrades gracefully when the extension is missing; lmdeploy/pytorch/__init__.py is empty aside from its header, meaning the actual PyTorch engine implementation lives outside the files read here. lmdeploy/__init__.py re-exports pipeline, serve, and client from a single api module alongside Pipeline, Tokenizer, GenerationConfig, TurbomindEngineConfig, PytorchEngineConfig, VisionConfig, and KVTransferConfig, and lmdeploy/serve/__init__.py separately exposes AsyncEngine, VLAsyncEngine, SessionManager, Session, and MultimodalProcessor for the serving path. setup.py builds the TurboMind extension via cmake_build_extension.CMakeExtension and pybind11 when LMDEPLOY_TARGET_DEVICE is cuda (the default) and DISABLE_TURBOMIND is not set, and separately resolves CUDA-version-pinned dependencies like nvidia-nccl-cu{CUDAVER} through get_turbomind_deps(). pyproject.toml scopes ruff linting to exclude third_party and src/turbomind from lint, implying the TurboMind C++ source sits outside the linted Python tree. docs/en/get_started/index.rst lists get-started guides for ascend, maca, and camb platforms, confirming non-Nvidia accelerator support beyond the default CUDA build path.

Observed

Packaging
Installable via pip from PyPI; the compiled engine additionally requires a cmake and pybind11 build toolchain and a CUDA-capable environment.
Engine options
Ships two distinct inference engines in one package: a compiled engine and a pure-Python engine, selected per run.
Interfaces
Exposes three interfaces: a command-line tool with subcommands, a Python library (pipeline, serve, client), and a documented API server.
Platform support
Includes get-started documentation for non-Nvidia accelerator platforms (Ascend, MetaX, Cambricon) alongside the default CUDA path.
Python support
Declared Python compatibility spans 3.10 through 3.13.
CLI options
The command-line chat subcommand exposes tensor-parallel, expert-parallel, prefix-caching, and speculative-decoding arguments directly.
Runtime checks
The compiled engine's absence is handled with a runtime availability check rather than an import failure.

Read from README.md, pyproject.toml, setup.py, lmdeploy/__init__.py, lmdeploy/cli/cli.py, lmdeploy/serve/__init__.py, lmdeploy/turbomind/__init__.py, lmdeploy/pytorch/__init__.py, docs/en/get_started/index.rst.

What it can do

  • Serve large language and vision-language models via an OpenAI/Anthropic-compatible API server

    LLM/VLM model → API endpoint for inference

  • Quantize model weights

    Model weights → Quantized weights

  • Quantize KV cache

    KV cache data → Quantized KV cache

  • Run inference using PyTorch engine

    Model and input prompt → Generated text/response

  • Run inference using TurboMind engine

    Model and input prompt → Generated text/response

  • Compress large language and vision-language models

    LLM/VLM model → Compressed model

Tags

llm-inferencequantizationmodel-servingturbomindvllm-alternativecudavision-language-models

Tech Stack

Python

Comments (0)

No comments yet

Editorially curated, with community endorsements as a secondary signal. Corrections welcome.