EmbeddingGemma 2 is an open-weight multimodal embedding model released by Google DeepMind on October 6, 2026 through the Google Developers Blog. Built on Gemma 4, it has just 740 million parameters, yet it maps text, images, video frames and audio natively into a single vector space, so a phone can run cross-modal retrieval locally instead of captioning images, transcribing speech and feeding separate models for each step.
Its small size comes from detachable encoders
It runs on-device because of a modular design: the text, vision and audio encoders load only when needed, so a photo-search feature does not carry the audio stack. On a Google Pixel 11 Pro, the text-only weights take roughly 191MB of memory and the full multimodal model about 567MB. The weights are compressed to INT4 and INT8 with quantization-aware training, output vectors are 768-dimensional, and Matryoshka Representation Learning lets developers truncate them on the fly to 128–512 dimensions, cutting local index storage by up to 8x with limited accuracy loss.
Search photos and find video moments offline
The official demos live in the Google AI Edge Gallery app. Instant Media Search re-ranks results as you type and can also search by example image; Video Moments Finder indexes local videos so a description like "kids laughing" jumps to the right timestamp without transcribing audio or generating captions first. The model also works as a zero-shot decision engine, matching inputs directly against candidate labels in milliseconds with no fine-tuning. On performance, a single image's visual embedding takes about 37.3 milliseconds on a MacBook M5 Pro GPU.
Three routes in for developers
For the easy path, MediaPipe Tasks wraps preprocessing and on-device vector search in the Embedder and Semantic Retriever, with one codebase across iOS, macOS, Windows, Linux and Web. For fine-grained control, LiteRT runs a single .litertlm file across CPUs, GPUs and NPUs. Android developers can also wait for the ML Kit integration arriving in the coming weeks, with NPU acceleration and automatic updates. The Mac app AI Edge Foresight, launched the same day, uses it to search local meeting records. The value of models like this is not a leaderboard: photo libraries and meeting notes are among the most sensitive data people own, and they can finally be searched without leaving the device.