EmbeddingGemma 2: One Open Embedding Model for Text, Code, Images, Video and Audio

| | 6 min read

Quick summary: EmbeddingGemma 2 is Google’s new open embedding model. It maps text, code, images, video and audio into one shared 768-dimensional vector space. The full model has 740M parameters (270M text, 170M vision, 300M audio) and ships under Apache 2.0 with an 8K context window. You can load only the text part for small deployments, and Unsloth publishes GGUF builds for running it locally.

Most retrieval stacks end up with a pile of separate models. One embeds your documents, another handles screenshots, and a third deals with audio transcripts. Each has its own vector space, so you cannot search one with the other. EmbeddingGemma 2 targets exactly that problem by putting every modality into a single space.

Google released it on October 6, 2026, and Unsloth announced GGUF builds and fine-tuning support the same day. Here is what it is, how it compares with the first EmbeddingGemma, and what to check before you adopt it.

What EmbeddingGemma 2 actually is

According to Google’s developer guide, EmbeddingGemma 2 is a modular multimodal embedding model built on Gemma 4. Instead of one monolithic network, it is a 270M text model with optional vision and audio encoders attached.

  • Text and code only: 270M parameters.
  • Text and vision: about 440M parameters.
  • Text and audio: about 570M parameters.
  • Everything: 740M parameters (270M text, 170M vision, 300M audio).

All inputs land in the same 768-dimensional space, and the context window is 8,192 tokens shared across modalities. The guide lists fixed token costs per input: one token per text token, 576 for an image, 16 per second of audio and 2,304 per second of video. It supports 100+ languages and visual documents such as PDFs, slides and charts. The license is Apache 2.0.

Why one embedding space matters for RAG

If you have built retrieval before, the value is easy to see. A single space means a text query can retrieve a slide, a chart, a clip or a voice note, with no captioning or transcription step in the middle. That removes a whole stage from many pipelines, and one less stage is one less place for errors to hide.

It also keeps your index simple. You store one vector type and one ANN index, whether that is an HNSW graph or something more compact such as vector quantisation. If you are still deciding whether embeddings are worth the trouble, my earlier piece on grep versus vector search covers when plain text search is enough.

EmbeddingGemma 2 benchmark results

The scores below come from the comparison graphic Unsloth published with the launch, and they roughly match Google’s model card. Treat them as vendor-reported numbers. Some competitor scores in that graphic are not self-reported by the competitor, so rerun on your own data before you decide.

Benchmark EmbeddingGemma 2 (740M) EmbeddingGemma 1 (308M) Jina v5 Omni-Nano Qwen3-Embedding 0.6B
MTEB multilingual v2 61.4 61.2 65.5 64.3
MTEB English v2 68.5 69.7 71.1 70.7
MTEB code v1 78.7 68.8 n/a 75.4
MIEB lite (image) 64.6 not supported 49.4 not supported
MMEB v2 VisDoc 67.8 not supported 71.4 not supported
MMEB v2 video 50.7 not supported 31.2 not supported
MSEB audio retrieval 69.5 not supported n/a not supported
MAEB audio 49.4 not supported 50.1 not supported

The honest reading is mixed. On pure text, EmbeddingGemma 2 does not beat the strongest small competitors, and its English score is slightly below the first EmbeddingGemma. Code retrieval is the clear text win. The real story is breadth: it is one of few open models with competitive numbers across image, video and audio at once.

Running EmbeddingGemma 2 locally with Unsloth GGUFs

Unsloth’s GGUF repository lists quantized builds from 4-bit (about 176 MB) up to 16-bit (about 558 MB). The repo shows this launch command for llama.cpp:

llama serve -hf unsloth/embeddinggemma-2-GGUF:UD-Q4_K_XL

On memory, Unsloth says the 270M text and code model runs in about 0.5 GB of RAM, and the full multimodal model needs about 1 GB. A Google edge post cites different on-device figures for a phone, so measure on your own hardware. Unsloth also says you can fine-tune the model through Unsloth Desktop, which I have not tested.

Three practical details from the model card

  • Use task prefixes. For retrieval, queries use task: search result | query: {query} and documents use title: {title} | text: {content}. Without them, text quality drops.
  • Avoid float16. Google says to use bfloat16 or float32, because float16 can cause silent degradation or NaN outputs. Prefer the BF16 file over F16 if you run the 16-bit build.
  • Truncate dimensions carefully. Matryoshka training lets you cut vectors to 512, 256 or 128 dimensions. Quality is close to lossless down to 256, but 128 hurts multimodal results noticeably. At 128 dimensions, a million vectors take roughly 250 MB instead of 1.5 GB.

What EmbeddingGemma 2 does not solve

The model card admits it struggles with sarcasm and figurative language, and quality varies across languages. The shared 8K context also means a long video can use up the whole window quickly. Embeddings find similar content, but they do not judge whether a retrieved chunk answers the question. If your RAG system fails on retrieval quality, a better embedder is only one fix, and techniques like contextual retrieval address a different part of the problem.

It is also worth separating this model from Gemini Embedding 2, a closed, API-only multimodal model with 3,072-dimensional vectors, described in its technical report. EmbeddingGemma 2 is the open, small, self-hostable option.

Should you try EmbeddingGemma 2?

Try it if you want one local index across documents, images and audio, or if you need a small code-aware embedder on modest hardware. Stay with a text-only model if all you embed is English prose and you already have a tuned pipeline, since the text benchmarks show no clear gain there.

A sensible first test: take 200 of your own queries, embed a real slice of your content with the 270M text part alone, then add the vision or audio encoder and compare recall. That tells you more than any leaderboard, and it shows whether the extra modalities are worth their memory.

Sources

The post EmbeddingGemma 2: One Open Embedding Model for Text, Code, Images, Video and Audio appeared first on Alpesh Kumar.

Subscribe to Our Newsletter

We don’t spam! Read our privacy policy for more info.