
Lets teams bootstrap speculative-decoding drafter models from existing pretrained small LMs instead of retraining from scratch for every target model, cutting the cost of speeding up deployments.
“This is a striking inversion of modern LLM development, where target models are valued precisely for the broad generalization they acquire through large-scale pretraining.”
“Osprey addresses both by pruning to a shallow backbone, restoring its language-modeling capability with target-agnostic next-token pretraining, and adapting it to each target through vocabulary alignment, zero-initialized QKV expansion, and distillation from the target model's output distribution.”
articleMargins, Not Windows: Training-Free Per-Step Lossy Speculative DecodingOszk\'ar Urb\'an, Young D. Kwon, Stylianos I. Venieris, Cecilia Mascolo
articleCo-Designing AI Models Using Speculative Decoding for Faster LLM InferenceTanya LenzChecking sign-in…
Loading comments…