EmbeddingGemma 2: an open, lightweight multimodal embedding model
Source
deepmind.google
Date
Why it matters
One 740M-parameter Apache 2.0 model embeds text, code, images, video and audio into a shared space. Builders can run cross-modal search and RAG on-device without a proprietary hosted embeddingA list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.Full definition → API.
Key takeaways · AI-distilled
The model is modular: Google says text-only workloads need as little as 270M parameters, with an optional 170M vision encoder and 300M audio encoder added only when you need full multimodal support.
Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256 or 128, which Google says cuts local vector-database storage and memory by up to 6x.
On a Pixel 11 Pro with quantizationShrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.Full definition →, Google reports roughly 191MB of active RAM for text-only weights and about 567MB for the full multimodal model.
The context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → is 8K tokens, four times EmbeddingGemma 1, enough per Google for up to 5.5 minutes of audio, 29 images or 58 video frames, or interleaved mixes of them, in one input.
Google reports code retrieval improved most: MTEB Code rose from 68.76 to 78.68 over EmbeddingGemma 1, while multilingual text performance held steady, pitching it for local codebase indexing and coding-AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → retrieval.
Terms in this piece · Glossary
embedding — A list of numbers representing a piece of text's meaning, so that similar meanings end up numerically close and can be searched.
quantization — Shrinking a model by storing its numbers less precisely — like rounding — so it runs faster and fits on smaller hardware, at a small quality cost.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.