NewsStocksGoogle DeepMind Launches EmbeddingGemma 2, an Open, Lightweight Multimodal Embedding Model

Google DeepMind Launches EmbeddingGemma 2, an Open, Lightweight Multimodal Embedding Model

Author: Google DeepMind Blog·

Key Takeaways

  • •Google DeepMind released EmbeddingGemma 2 on October 6, 2026, an open 740-million-parameter multimodal embedding model built on the Gemma 4 architecture and licensed under Apache 2.0.
  • •The model natively unifies text, code, images, video, and audio in a single embedding space, and posts leading scores among sub-1B multimodal embedders on benchmarks such as MTEB Code and MAEB, with its code score rising 9.92 points to 78.68.
  • •Its modular design requires as little as 270M parameters for text-only workloads, with optional vision and audio encoders, and supports truncating vectors from 768 to as low as 128 dimensions for up to 6x storage reduction.
  • •With quantization, the model needs roughly 191MB of active RAM for text-only weights and about 567MB for the full multimodal version on a Google Pixel 11 Pro, and offers an 8K token context window four times larger than its predecessor.
  • •Generating embeddings locally supports privacy-preserving, offline semantic search and on-device RAG when paired with Gemma 4, with model weights available on Hugging Face and Kaggle and compatibility across a wide range of developer tools.
Google DeepMind Launches EmbeddingGemma 2, an Open, Lightweight Multimodal Embedding Model

Google DeepMind on October 6, 2026 launched EmbeddingGemma 2, an open embedding model that the company describes as the most capable to date for on-device multimodal embeddings, natively mapping combinations of text, images, audio, and video into a unified embedding space. Embedding models compress content into numeric vectors whose relative positions encode meaning, and they serve as the retrieval layer in semantic search and retrieval pipelines. The announcement was published on the Google DeepMind blog by research engineers Sahil Dua and Henrique Schechter Vera.

Google DeepMind introduced the original EmbeddingGemma last year as a lightweight option for high-quality text embeddings, designed to help applications organize, search, and connect information directly on consumer hardware. According to the company, the developer community's response exceeded expectations: the model has been downloaded more than 20 million times, and builders have used it to power smarter on-device search tools and privacy-first retrieval augmented generation (RAG) pipelines.

EmbeddingGemma 2 extends the approach beyond text, unifying code, images, video, and audio in a shared embedding space. The model is built on the Gemma 4 architecture, released under the commercially permissive Apache 2.0 license, which allows commercial use, modification, and redistribution without licensing fees, and carries 740 million parameters, a size Google DeepMind says is optimal for on-device inference. Possible uses include locating a specific video clip from a voice memo or searching through hours of audio recordings with a text query, all processed by a single, natively multimodal model.

Built from the same technology as Gemini Embedding models, EmbeddingGemma 2 has the following characteristics, according to the announcement:

  • Best-in-class for its size: Leading scores among sub-1B multimodal embedders for its size across benchmarks such as MTEB (Massive Text Embedding Benchmark) Code and MAEB (Massive Audio Embedding Benchmark), while matching or outperforming many larger models across text, vision, and audio tasks.
  • Modular by design: Requires as little as 270M parameters for text-only workloads, with optional vision (170M) and audio (300M) encoders for full multimodal support.
  • Storage-efficient: Using Matryoshka Representation Learning (MRL), developers can dynamically truncate output vectors from 768 dimensions down to 512, 256, or 128 dimensions, providing up to 6x storage reduction for local vector databases and memory usage.
  • Optimized for on-device performance: Runs efficiently within tight resource constraints. With quantization, on a Google Pixel 11 Pro the model requires as little as ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model.
  • Extended context ready: Features an 8K token context window, four times larger than EmbeddingGemma 1, allowing it to process up to 5.5 minutes of audio, 29 images, 58 video frames, or interleaved combinations thereof directly on local hardware.

Top-tier quality for code, vision, and audio

EmbeddingGemma 2 matches the strong multilingual text performance of its predecessor while delivering a 9.92-point improvement on code performance in MTEB Code, rising from 68.76 to 78.68. Google DeepMind says this makes the model well-suited for local codebase indexing, semantic code search, and coding agent retrieval. Across image, video, documents, and audio, the company states that the model sets a new standard in quality-per-parameter for sub-1B models and even outperforms some specialist models more than twice its size. Full evaluation metrics and model information are available in the EmbeddingGemma 2 model card.

Semantic search, routing, and retrieval on-device

Google DeepMind positions EmbeddingGemma 2 as bringing search, routing, and retrieval capabilities directly to edge hardware. Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and allows developers to build cross-modal search and retrieval that works entirely offline. For applications handling sensitive media, local inference keeps the underlying content on the device even as search and retrieval operate across it.

When paired with generative models such as Gemma 4, EmbeddingGemma 2 enables on-device RAG pipelines that understand complex multimodal data. Because the model is built on Gemma 4 and shares its text tokenizer and audio encoder, developers can run both models together in a unified pipeline with a lower combined total memory footprint.

The company demonstrated several applications: using text or an image to top matches in a media library based on semantic similarity via Google AI Edge Gallery's Instant Media Search; locating specific moments in video with text or audio queries through the Gallery's Video Moments Finder; pairing EmbeddingGemma 2 for local file retrieval with Gemma 4 for contextual reasoning in the Google AI Edge Foresight app; and creating real-time decision engines that leverage multimodal context for classification, routing, and predictive capabilities via the MediaPipe Decision Task API. A Google AI Edge blog post explains how to build on-device search and RAG systems with LiteRT.

Getting started with EmbeddingGemma 2

Google DeepMind said it worked closely with partners to ensure the model works immediately in common developer environments. Model weights are available on Hugging Face and Kaggle, with availability on the Gemini Enterprise Agent Platform Model Garden coming soon. LiteRT Community on Hugging Face hosts models optimized for on-device use.

For deployment, developers can build cross-platform apps with Google AI Edge MediaPipe for turnkey embedding, retrieval, and decision tasks, or use LiteRT for custom model integration, and build for the browser with transformers.js or WebGPU.

The model can be served efficiently with tools including transformers, sentence-transformers, MLX, vLLM, llama.cpp, SGLang, Ollama, and LMStudio, with embedding vectors storable in Qdrant. Guidance from Unsloth covers fine-tuning the model for specific use cases. A developer guide, documentation, and guides for inference and fine-tuning are also available.