Google releases EmbeddingGemma 2 for on-device multimodal search
The 740-million-parameter model maps text, code, images, audio and video into a shared space. Google says it can run on a phone or small board under an Apache 2.0 licence.
Google says EmbeddingGemma 2 has 740 million parameters and can run entirely on a phone or a small board. The model maps text, code, images, audio and video into a shared space. Its release makes an open embedding model available for on-device retrieval, while leaving performance differences and the absence of safety tuning as relevant limits.
Embedding models turn content into numerical representations that software can compare by meaning. Google says the new model uses a 768-dimension space for all five modalities. It is licensed under Apache 2.0, a licence that permits broad use and modification subject to its terms.
On-device use keeps content local
Google says the model can perform text work using about 191MB of memory on a Pixel 11 Pro. Loading all five modalities requires about 567MB, according to the source material. The account says the retrieval process need not send content off the device, so the intended design keeps that material local.
The local approach may matter when developers handle sensitive conversations or other material they do not want transferred. The source describes avoiding network transfer as a practical feature of an on-device retrieval pipeline. It does not establish how any particular deployment will handle retained data, access controls or user consent.
The model separates modality components
Of the 740 million parameters, Google says 270 million are needed for text alone. A 170-million vision encoder and a 300-million audio encoder can be loaded when wanted. The model also shares its text tokeniser and audio encoder with Gemma, according to the source, which says the two can run together in less combined memory.
Google says EmbeddingGemma 2 supports more than 100 languages, but performance may not be equal across them. The source also says the model has received no safety tuning and has no output moderation. Mitigation was applied to training data instead, so those limits form part of the model’s operating context.
Google has not described equal coverage
The available information does not give language-by-language performance results or identify which languages have weaker results. It also does not describe a regulator’s assessment, a legal proceeding or a formal regulatory stage. There is no company response to report on a dispute or investigation, because none is identified in the source material.
The account says the model supports 8,192 tokens across modalities, and Google says vectors can be reduced from 768 dimensions to 128. It gives examples of capacity as 29 images, 58 video frames or five and a half minutes of audio. The source does not explain the conditions behind each capacity example or how reduction affects retrieval quality.
Google’s stated position is that the model can support retrieval without sending content over a network. The next useful evidence would include comparative results across supported languages and details about performance after vector reduction. Developers will also need to assess the lack of safety tuning and output moderation against their own uses. The source provides no date for a regulatory review or a company response to one.