RELEASE · MODELS · #1603
Google releases EmbeddingGemma 2 — open 740M multimodal on-device embedding model
EmbeddingGemma 2 is a 740M-parameter, Apache 2.0–licensed multimodal embedding model from Google that natively maps text, images, audio, video and code into a single embedding space and is optimized for on-device use. It is built on the Gemma 4 architecture, supports modular encoder configurations, an 8K-token context window, Matryoshka Representation Learning for configurable vector sizes, and is available on Hugging Face and Kaggle.
KEY POINTS
- EmbeddingGemma 2 is a 740M-parameter, Apache 2.0–licensed multimodal embedding model from Google that natively maps text, images, audio, video and code into a single embedding space and is optimized for on-device use.
- It is built on the Gemma 4 architecture, supports modular encoder configurations, an 8K-token context window, Matryoshka Representation Learning for configurable vector sizes, and is available on Hugging Face and Kaggle.
- EmbeddingGemma 2 makes high-quality multimodal retrieval and on-device RAG feasible with sub-1B sizes, improving privacy and latency for offline, cross-modal search and indexing workflows.
WHY IT MATTERS
EmbeddingGemma 2 makes high-quality multimodal retrieval and on-device RAG feasible with sub-1B sizes, improving privacy and latency for offline, cross-modal search and indexing workflows.
SOURCES & TIMELINE
3EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space. We introduced EmbeddingGemma last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardware. The developer community’s response blew past ou…
EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space. Your browser does not support the audio element. We introduced EmbeddingGemma last year to provide a lightweight option for high-quality text embeddings, to help your apps organize, search, and connect information directly on consumer hardwar…
Google released EmbeddingGemma 2, an open model that converts text, images, video, audio, and code into numerical vectors so similar content can be found and compared more easily. At 740 million parameters, Google says it's the most compact model of its kind and outperforms competing models up to twice its size on multimodal embedding benchmarks. The model runs locally without an API key. Each query takes about 20 t…