Google launches EmbeddingGemma 2 for on-device multimodal retrieval
The model is built for local search and retrieval, letting more data stay on device.
Google has released EmbeddingGemma 2, a 740 million parameter multimodal embedding model aimed at on-device inference for text, code, images, video, and audio. The company says it is built on Gemma 4, uses an Apache 2.0 license, and can support local search and retrieval workflows without sending data off device. Google also claims a 9.92-point gain on code benchmarks versus the previous EmbeddingGemma, with MTEB Code rising from 68.76 to 78.68. The model supports an 8K token context window and can be paired with Gemma 4 for on-device multimodal RAG pipelines with lower combined memory use.
Why it matters
For users and teams handling text, code, images, video or audio, the main change is that embedding-based search and retrieval can run locally rather than relying on a remote service. Google also says the new model improves code performance over the previous EmbeddingGemma and supports an 8K token context window, which broadens the kinds of local workflows it can support.
Keep or strike?
Does this story matter, or is it hype? Mark it before you see what everyone else did.
Sources
- Google Blog