InferenceX DeepSeek v4 Performance Tracker
newsletter.semianalysis.com- Category
- Developer Tools
- Pricing
- Open Source
- Type
- ARTICLE
- Added
- Aug 2, 2026
About
An open-source SemiAnalysis initiative that benchmarks DeepSeek v4 Pro inference performance across hardware (GB300 NVL72, Huawei Ascend 950DT, MI355X, B200, H200) and engines (vLLM, SGLang, TensorRT-LLM, ATOM) from Day 0 through subsequent weeks, tracking how throughput and interactivity improve as vendors patch and optimize their stacks. The article documents specific bugs (e.g., a hardcoded hidden-size guard in TensorRT-LLM, broken AITER kernels on AMD ROCm) and shows AMD's MI355X achieving over 100x throughput improvement within 26 days via kernel-level fixes.
What it can do
Benchmark DeepSeek v4 Pro inference performance across different hardware and inference engines
Hardware/engine configuration (e.g., GB300 NVL72, MI355X, vLLM, TensorRT-LLM) → Throughput and interactivity performance data
Why it made the leaderboard
If you're deciding which inference engine or GPU to deploy DeepSeek v4 on, this tracks real Day-0-to-week-4 performance evolution with specific bug reports and kernel fixes across vLLM, SGLang, TensorRT-LLM, and hardware from NVIDIA, AMD, and Huawei — letting you time your deployment around when a stack actually becomes usable rather than trusting vendor marketing.
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.