EmbeddingGemma 2: an open, lightweight multimodal embedding model
Source
blog.google
Date
Why it matters
One small open model can now index mixed media on-device in a shared space, enabling multimodal and private RAG without hosted embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition →. Apache 2.0 permits commercial use.
Key takeaways · AI-distilled
The model is modular: text-only workloads need as little as 270M parameters, with optional vision (170M) and audio (300M) encoders added for full multimodalA model that works with more than text — reading images, audio, or video, and sometimes generating them too.Full definition → support.
Google says that with quantizationShrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.Full definition → on a Pixel 11 Pro it needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
Matryoshka Representation Learning lets developers truncate vectors from 768 dimensions to 512, 256, or 128, cutting local vector storage and memory by up to 6x.
The context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → is 8K tokens, four times the first EmbeddingGemma, enough for about 5.5 minutes of audio, 29 images, or 58 video frames in one input.
Code retrieval improved most: MTEB Code rose from 68.76 to 78.68. Because it shares Gemma 4's text tokenizer and audio encoder, running both together in one pipeline lowers combined memory.
Terms in this piece · Glossary
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
multimodal — A model that works with more than text — reading images, audio, or video, and sometimes generating them too.
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.