Vibeleaderboard
← All Intel
Intel / article

Giga-Embeddings: Mixture-of-Experts Encoders for High-Throughput Text Embeddings

Source
arxiv.org
Author
Egor Kolodin, Egor Krasnoperov, Evgeniy Kosarev, Fyodor Minkin
Date
Why it matters

A sparse encoder delivers 25% more throughput than a dense 3B at stronger retrieval quality, with all three checkpoints released — a real option for high-volume indexing pipelines.

Terms in this piece · Glossary
  • embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
  • mixture-of-experts — A model built from many specialist sub-networks where only a few activate per token, giving big-model capability at small-model running cost.
  • token — The chunk of text a model reads and writes in — roughly three-quarters of a word — and the unit AI usage is billed in.
  • distillation — Training a small, cheap model to imitate a big one's outputs, keeping much of the capability at a fraction of the cost.
Recommended reads
Comments

Checking sign-in…

Loading comments…