Vibeleaderboard
← All Intel
Intel / blog

EmbeddingGemma 2: an open, lightweight multimodal embedding model

Source
blog.google
Date
Why it matters

One small open model can now index mixed media on-device in a shared space, enabling multimodal and private RAG without hosted . Apache 2.0 permits commercial use.

Key takeaways · AI-distilled
  • The model is modular: text-only workloads need as little as 270M parameters, with optional vision (170M) and audio (300M) encoders added for full support.
  • Google says that with on a Pixel 11 Pro it needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
  • Matryoshka Representation Learning lets developers truncate vectors from 768 dimensions to 512, 256, or 128, cutting local vector storage and memory by up to 6x.
  • The is 8K tokens, four times the first EmbeddingGemma, enough for about 5.5 minutes of audio, 29 images, or 58 video frames in one input.
  • Code retrieval improved most: MTEB Code rose from 68.76 to 78.68. Because it shares Gemma 4's text tokenizer and audio encoder, running both together in one pipeline lowers combined memory.
Terms in this piece · Glossary
  • embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
  • multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
  • quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
  • context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Read the source blog.google
Recommended reads
Comments

Checking sign-in…

Loading comments…