Vibeleaderboard
Index / tool
Visit dflash.z-lab.ai
Category
AI Tools
Pricing
Open Source
Type
TOOL
Builder
z-lab
Added
Jun 18, 2026

About

DFlash is a lightweight block diffusion model designed for speculative decoding that accelerates LLM inference through efficient parallel token drafting. It provides pre-trained draft models for popular LLMs like Qwen and LLaMA, enabling significant speedups in text generation.

Why it made the leaderboard

A lightweight block diffusion model for speculative decoding: it drafts tokens in parallel to accelerate LLM inference, with pre-trained draft models for Qwen and LLaMA. Works across vLLM, SGLang, Transformers, and MLX, so it slots into an existing serving stack.

Tags

llminferencespeculative-decodingoptimizationtransformersvllmsglang

Tech Stack

Python

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.