EmbeddingGemma 2: An Open, Lightweight Multimodal Embedding Model
AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get tech for your team delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Google announced EmbeddingGemma 2 on Oct. 6, 2026, an open-weight, 740-million-parameter model that maps text, code, images, audio and video into a shared embedding space. Google says it is designed for local use, supports an 8,192-token context window and is available under the Apache 2.0 license; independent verification of its performance claims is not included in the announcement.

Google announced EmbeddingGemma 2 on Oct. 6, describing it as an open, 740-million-parameter model that turns text, code, images, audio and video into vectors in a shared embedding space. The release expands the EmbeddingGemma line beyond text and is aimed at developers building search and retrieval features that can run on local devices rather than sending content to a server.

Google says the model is based on the Gemma 4 architecture and is released under the commercially permissive Apache 2.0 license. Its multimodal design is intended to support cross-format searches, such as locating a video clip using a voice memo or finding relevant audio recordings with a text query. Google says these tasks can be handled by one model on-device.

The model has an 8,192-token context window, which Google says can cover up to 5.5 minutes of audio, 29 images or 58 video frames, as well as combinations of those inputs. Developers can shorten output vectors from 768 dimensions to 512, 256 or 128 using Matryoshka Representation Learning. Google says this can reduce vector storage and memory use by up to six times, depending on the selected size and workload.

Google also describes a modular configuration: text-only use can require as little as 270 million parameters, with optional vision and audio encoders listed at 170 million and 300 million parameters, respectively. For a quantized deployment on a Google Pixel 11 Pro, the company reports active RAM use of about 191 MB for text-only weights and 567 MB for the full multimodal model. Those figures are Google’s stated results and may vary with implementation and hardware.

At a glance
announcementWhen: Announced Oct. 6, 2026
The developmentGoogle announced EmbeddingGemma 2, a multimodal embedding model intended to run on consumer hardware and support local search and retrieval.

Local Search Across Media Types

EmbeddingGemma 2 targets a practical limitation in local search: information may be spread across notes, recordings, images and video, while many search systems are built around one format at a time. A shared embedding space could let an app retrieve related material across those formats, including when the search query and result use different media.

Running embedding generation locally may also reduce the need to upload personal content to a remote service and can allow retrieval to work without an internet connection. That could matter for private archives, field work and apps with limited connectivity. These are potential uses, not guarantees: privacy depends on how developers design the rest of an app, and the announcement does not provide independent tests of end-to-end products.

The model is also relevant to developers working within device memory limits. Google’s stated RAM figures and adjustable vector sizes offer options for tailoring deployments, though the trade-off between storage savings, speed and retrieval quality will depend on each application. The Apache 2.0 release permits commercial use under its terms, giving teams more room to inspect and adapt the model than a closed hosted service would.

Amazon

on-device multimodal AI model

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Text Embeddings to Multimodal

Google introduced the original EmbeddingGemma as a lightweight text-embedding model for tasks such as organizing information and building search and retrieval-augmented generation systems. In its Oct. 6 announcement, Google said that first model had received more than 20 million downloads. The company did not provide a date range or independent download count in the material supplied.

EmbeddingGemma 2 keeps the focus on smaller deployments while adding image, audio and video inputs. Google says it matches the first model’s multilingual text performance and improves its MTEB Code score from 68.76 to 78.68, a 9.92-point increase. The company also reports leading results among sub-one-billion-parameter multimodal embedding models on selected benchmarks, including MTEB Code and the Massive Audio Embedding Benchmark. Full evaluation details are directed to the model card; benchmark claims here are attributed to Google.

Google lists model weights on Hugging Face and Kaggle. Its announcement also points developers to Google AI Edge MediaPipe and LiteRT for on-device deployment, and to browser and model-serving tools including transformers.js, MLX, vLLM and llama.cpp. Availability through Gemini Enterprise Agent Platform Model Garden was described as coming soon, not as a current release.

““EmbeddingGemma 2 is the most capable model for on-device multimodal embeddings.””

— Sahil Dua and Henrique Schechter Vera, Google DeepMind research engineers, in Google’s announcement

Amazon

portable AI embedding device

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Performance and Device Trade-Offs

The announcement does not include independent evaluations of the model’s benchmark results, real-world retrieval quality or performance across different devices. Google directs readers to the model card for full metrics, but the supplied material does not give details such as test conditions, confidence intervals or comparisons for every task. The company’s description of EmbeddingGemma 2 as the most capable in its category should be treated as a vendor claim.

It is also not yet clear how much the model’s stated memory requirements and sixfold storage reduction affect speed or accuracy in specific applications. The RAM numbers are for a quantized Pixel 11 Pro setup, and performance may differ across hardware and software configurations. The announcement does not detail the model’s limitations with particular languages, media formats or difficult retrieval queries.

Amazon

multimodal search software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Model Access and Developer Testing

Developers can download the announced weights from Hugging Face and Kaggle and test deployments through the listed Google AI Edge tools and other supported frameworks. Google’s model card and developer documentation are the next sources for checking evaluation methods, inference requirements and fine-tuning guidance.

Google said availability in Gemini Enterprise Agent Platform Model Garden is coming soon, without giving a date. Wider developer testing should help establish how the model performs on specific devices and whether its memory, latency and retrieval-quality trade-offs hold across different applications.

Amazon

AI-powered multimedia search tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is EmbeddingGemma 2?

It is a 740-million-parameter embedding model announced by Google for mapping text, code, images, audio and video into a shared vector space for search and retrieval.

Can EmbeddingGemma 2 run without an internet connection?

Google says it is designed for on-device inference and that cross-modal retrieval can work offline. Whether an app works fully offline depends on how developers build and deploy it.

What license does Google use?

Google says the model is released under the Apache 2.0 license, which permits commercial use subject to the license terms.

Where can developers get the model?

Google lists model weights on Hugging Face and Kaggle. The company said availability in Gemini Enterprise Agent Platform Model Garden was coming soon.

How much memory does it use?

For a quantized setup on a Google Pixel 11 Pro, Google reports about 191 MB of active RAM for text-only weights and 567 MB for the full multimodal model. Actual requirements can vary by device and deployment.

Source: hn

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

$100 AI Music Video: Claude Fable 5 Vs. GPT-5.6 Sol

Two AI models, Claude Fable 5 and GPT-5.6 Sol, compete in creating a music video for $100, highlighting advances in AI-generated content.

U.S. Lifts Restrictions on Anthropic’s Most Powerful A.I. Models

The U.S. government has removed restrictions on Anthropic’s most advanced AI models, enabling wider deployment and use in various sectors.

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Analyzing Mistral’s focus on European sovereignty, open weights, and local infrastructure amid Europe’s AI ambitions and global competition.

Scriptc By Vercel: TypeScript-to-Native Compiler, No JavaScript Engine In Binary

Vercel introduces Scriptc, a TypeScript-to-native compiler that produces binaries without embedding a JavaScript engine, streamlining deployment.