
🚀 We’ve open-sourced AngelSpec, an end-to-end speculative decoding framework supporting both training and deployment. On Hy3-A21B, DFly delivers a 1.98–2.40× end-to-end speedup over autoregressive decoding across tested concurrency levels from 4 to 64, with 10.5–11.8% higher throughput than DFlash. Training code and Hy3-A21B MTP/DFly drafter weights are now available: GitHub: https://t.co/6Li9K4wTFY Paper: https://t.co/BxJf0Gr3Nv Docs: https://t.co/BeFFxxsGKR Hugging Face: https://t.co/lT5tWoto1C ModelScope: https://t.co/1LNTI906EY #Hy3 #AngelSpec #OpenSource
is one of the few wins needing no model change; here both the drafter training code and Hy3 drafter weights ship, with numbers across concurrency 4–64 rather than one best case.
Checking sign-in…
Loading comments…