
Hy4 preview runs in vLLM https://t.co/16PHpD9Fxm
You can serve Hy4-preview on vLLM today without waiting for a port, and the disclosed layout of 2048 attended per query with sparse index reuse across 57 of 78 layers explains how the 1M window stays affordable.
articleOne GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless inAlibaba_Qwen
articleRecent Developments in LLM Architectures: KV Sharing, mHC, and Compressed AttentionSebastian Raschka, PhD
articleEfficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon et al.Checking sign-in…
Loading comments…