
Large multilingual vocabularies make the output head a memory bandwidth bottleneck on CPU. Swapping the dense projection for approximate retrieval raised single-stream decoding throughput up to 82% on Gemma 3 270M with no measured quality loss.
articleRecipes for Steering and Scaling LLMs via SamplingJiajun He, Zongyu Guo, Jos\'e Miguel Hern\'andez-Lobato, Yuanqi Du
articleGiga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text EmbeddingsEgor Kolodin, Egor Krasnoperov, Evgeniy Kosarev, Fyodor Minkin
articleRENDER: Controlling Reader-Facing Evidence in LLM Memory EvaluationYuan Si, Simeng Han, Daming Li, Jialu ZhangChecking sign-in…
Loading comments…