Trusence Every claim has a source
Last updated 6 October 2026 Search Türkçe
← All stories
Models

Google launches EmbeddingGemma 2 for on-device multimodal retrieval

The model is built for local search and retrieval, letting more data stay on device.

Google has released EmbeddingGemma 2, a 740 million parameter multimodal embedding model aimed at on-device inference for text, code, images, video, and audio. The company says it is built on Gemma 4, uses an Apache 2.0 license, and can support local search and retrieval workflows without sending data off device. Google also claims a 9.92-point gain on code benchmarks versus the previous EmbeddingGemma, with MTEB Code rising from 68.76 to 78.68. The model supports an 8K token context window and can be paired with Gemma 4 for on-device multimodal RAG pipelines with lower combined memory use.

Why it matters

For users and teams handling text, code, images, video or audio, the main change is that embedding-based search and retrieval can run locally rather than relying on a remote service. Google also says the new model improves code performance over the previous EmbeddingGemma and supports an 8K token context window, which broadens the kinds of local workflows it can support.

Sources

  • Google Blog