- Google has released “EmbeddingGemma 2,” an open, multimodal embedding AI model designed to run directly on-device.
- EmbeddingGemma 2 now handles text, code, images, video, and audio together within a single model.
- It enables processing entirely on-device without an internet connection, such as finding specific video clips using voice memos or searching hours of audio recordings with text.
On Tuesday, October 6, 2026 (local time), Google’s AI research division, Google DeepMind, released “EmbeddingGemma 2,” an open, multimodal embedding AI model that operates on-device.
Embedding AI models convert data such as text and images into numerical values that represent their meaning. They serve as the foundation for searches and classification to find semantically similar data, as well as for RAG (Retrieval-Augmented Generation), which allows generative AI to reference local information to formulate answers.
While the first-generation “EmbeddingGemma,” launched in 2025, was text-only, the newly released “EmbeddingGemma 2” integrates text, code, images, video, and audio into a single model. This enables on-device execution entirely without an internet connection—such as finding a specific video clip using a voice memo or searching through hours of audio recordings using text.
“EmbeddingGemma 2” is based on Google’s open model “Gemma 4” and features 740 million parameters. For text-only use cases, it operates on 270 million parameters, with processing modules for images (170 million) and audio (300 million) added as needed. When quantized, memory usage on the “Pixel 11 Pro” is kept to approximately 191 MB for text-only, and about 567 MB even when utilizing all features.
It can process 8K tokens at a time—four times the capacity of the first generation—handling up to 5.5 minutes of audio, 29 images, 58 frames of video, or a combination thereof. Additionally, the output data size can be progressively reduced, cutting the storage footprint of search data on the device by up to six times.
In terms of performance, while maintaining the same multilingual text capabilities as its predecessor, its score on “MTEB Code,” a benchmark measuring code search performance, improved by 9.92 points from 68.76 to 78.68. It also boasts top-tier performance for models under 1 billion parameters across image, video, and audio tasks, outperforming specialized models more than twice its size in certain scenarios.
Furthermore, because it shares components with “Gemma 4,” developers can build fully on-device RAG systems with low memory consumption even when using the two models in combination. Google’s on-device AI experience app, “Google AI Edge Gallery,” has also been updated with features powered by “EmbeddingGemma 2”: “Instant Media Search,” which lets users search their photo library using text or images, and “Video Moments Finder,” which allows them to find specific scenes within videos using text or voice.
“EmbeddingGemma 2” is available on Hugging Face and Kaggle under the commercial-friendly Apache 2.0 license. It also supports major development tools such as transformers, Ollama, llama.cpp, and LM Studio.
Source: Google (https://blog.google/innovation-and-ai/technology/developers-tools/embeddinggemma-2/)





