Google brings multimodal semantic search to the edge with EmbeddingGemma 2
Google DeepMind has launched EmbeddingGemma 2, an open-weight multimodal embedding model built for local retrieval and decision-making. The 740-million-parameter model maps text, images, video frames and audio into one shared vector space. That lets a developer compare different media types directly, instead of chaining separate captioning, speech-to-text and text-embedding systems. Google says the full multimodal model can run with about 567 MB of active RAM on a Pixel 11 Pro, while text-only weights use about 191 MB. These are vendor-reported figures, not an independent device test.
The release is accompanied by two demos in the Google AI Edge Gallery app. Instant Media Search lets users find local photos and videos with natural-language queries, example images or a camera stream. Video Moments Finder indexes video keyframes and audio locally, then returns timestamps for prompts such as a dog catching a frisbee or children laughing. Google says the demos work without an internet connection and use local storage and similarity search.
Google also introduced AI Edge Foresight for Mac, an experimental meeting companion that indexes transcripts and private files on the device. The company says Foresight uses EmbeddingGemma 2 and Gemma 4 for note-taking and cross-modal retrieval. For Android developers, support through ML Kit is planned in the coming weeks. MediaPipe Tasks already provides a cross-platform route through its Universal Embedder and Semantic Retriever, while LiteRT offers lower-level control over CPU, GPU and NPU acceleration.
The practical importance is privacy and latency. A local embedding index can keep sensitive media and search queries on a phone or computer and can respond while someone is typing, even when connectivity is poor. That does not remove the need for device security, permission controls or careful retention policies. Local processing also shifts costs toward hardware, model packaging and index management.
For makers and businesses, EmbeddingGemma 2 is therefore a useful building block rather than a complete product. Its open weights and ready-made edge integrations may shorten the path to private media search, document retrieval and intent routing. Independent measurements, supported hardware and real-world accuracy will determine how broadly it fits production workloads.