Google Launches Compact EmbeddingGemma 2 Model
Google has launched EmbeddingGemma 2, a highly efficient open model that brings powerful multimodal embedding capabilities directly to local devices with minimal memory requirements.

Google has released EmbeddingGemma 2, a new open-weights model designed to convert text, images, video, audio, and code into numerical vectors. At 740 million parameters, the tech giant describes it as the most compact model of its kind. Despite its small size, Google claims the model outperforms competing embedding models up to twice its size on multimodal benchmarks, making high-quality vector search more accessible.
On the Massive Text Embedding Benchmark (Code), EmbeddingGemma 2 achieved a score of 78.68. This represents a significant jump of nearly 10 points over its predecessor, which scored 68.76. This leap in performance puts the compact model on par with much larger alternatives. For developers who only require text-processing capabilities, Google is also offering a smaller 270-million-parameter version of the model.
For practitioners, the model's efficiency is its primary advantage. EmbeddingGemma 2 runs entirely locally without requiring an API key, executing queries in just 20 to 70 milliseconds via WebGPU in a standard web browser. It requires only about 191 MB of RAM to operate and reduces local vector database storage requirements by up to six times, drastically lowering the hardware barrier for local deployment.
These resource savings allow developers to build private, offline applications. When paired with small open models like Gemma 4, EmbeddingGemma 2 can power offline retrieval-augmented generation systems that process sensitive data without transmitting it to external cloud servers. The model weights, along with developer guides and documentation, are currently available for download on Hugging Face and Kaggle.
This is our own summary of reporting by The Decoder



