EmbeddingGemma 2: how to generate embeddings locally without paying for an API

EmbeddingGemma 2 is Google’s new open embedding model: it runs locally and converts text, code, images, audio, and video into vectors, without paying per token to an API. It has 740M parameters, is published under the Apache 2.0 license, and its biggest improvement over the first version is in code search.

What is EmbeddingGemma 2 and what are embeddings?

An embedding is a numerical vector that represents the meaning of content: two texts with similar meaning produce nearby vectors, even if they don’t share words. It’s the piece that makes semantic search, classification, and any RAG (Retrieval-Augmented Generation) pipeline work.

EmbeddingGemma 2 is the second generation of Google DeepMind’s lightweight embedding model, built on the Gemma 4 architecture. It generates 768-dimensional vectors in a single shared space for four modalities: text (including code), images, video, and audio. Its components are:

  • Text model: 270M parameters (a 130M backbone plus a 140M embedder).
  • Vision encoder: 170M, optional.
  • Audio encoder: 300M, optional.

You only load the encoders you need, so a text and code pipeline never pays for vision or audio.

The context window is 8,192 tokens, four times that of the first version. Multimodal inputs share that budget: about 29 images, about 58 video frames, or around five and a half minutes of audio. Google indicates support for over 100 languages.

According to Google, the first version of EmbeddingGemma, released last year, surpassed 20 million downloads.

Is EmbeddingGemma 2 better for code search?

Yes, and code is what changed the most. On MTEB Code, EmbeddingGemma 2 scores 78.68 versus 68.76 for the previous version: 9.92 points higher, roughly 14% relative.

On multilingual text, the two versions are practically tied (61.36 versus 61.15). If you already use EmbeddingGemma 1 for RAG on plain text, the reason to upgrade is multimodality or code, not prose.

All these figures come from Google’s model card and use the checkpoint in full precision. They are self-reported results, and “best-in-class among multimodal embedders under 1B” is a category Google itself defines. Take them as a starting point and measure on your own repository.

Google positions the code improvement for three specific uses:

  • Index a codebase locally (local codebase indexing)
  • Semantic code search (semantic code search)
  • Context retrieval for code agents (code embeddings for coding agents)

That’s the realistic use case for most: give your agent or internal search engine a local index of your repository, without sending the code to a third party.

How to generate embeddings locally with Sentence Transformers?

The fastest way is Sentence Transformers. Install it along with Transformers:

pip install -U sentence-transformers transformers

Load the model in text-only mode with a safe numeric type:

import torch
from sentence_transformers import SentenceTransformer

dtype = torch.bfloat16 if torch.cuda.is_bf16_supported() else torch.float32

model = SentenceTransformer(
    "google/embeddinggemma-2",
    model_kwargs={"torch_dtype": dtype},
    config_kwargs={"vision_config": None, "audio_config": None},
)

Both entries in config_kwargs matter. By default, SentenceTransformer loads all 740M parameters, including vision and audio. By disabling both, you’re left with the 270M text model.

Now index some code and search it in natural language:

files = {
    "auth/session.py": "def refresh_token(user): ...",
    "billing/invoice.py": "def generate_invoice(order): ...",
}

doc_embs = model.encode(
    [f"title: {name} | text: {code}" for name, code in files.items()]
)
query_emb = model.encode("where is the session token refreshed?",
                         prompt_name="CodeRetrieval")

print(model.similarity(query_emb, doc_embs))

Two details come straight from the model card:

  • Task prefixes: the model expects prefixes in text. prompt_name="CodeRetrieval" applies task: code retrieval | query: to the query.
  • Documents with title: formatted by hand as title: {file} | text: {code}. Including the filename as the title improves retrieval.

The model card also lists other prompts:

  • SearchQuery / Document: general search
  • QuestionAnswering
  • Classification
  • Clustering
  • SentenceSimilarity

How to use EmbeddingGemma 2 with Ollama?

The model is already in Ollama’s library, so you can generate embeddings with Ollama (ollama embeddings) without writing Python:

ollama pull embeddinggemma-2

At the time of publishing this note, Ollama listed, among others, the tags :270m (378 MB) and :740m (1.3 GB). The metadata on that card doesn’t match Google’s model card on context window, so use the model card as reference for the limits.

If you already run Gemma 4 locally with Ollama, pairing them makes sense: EmbeddingGemma 2 shares with Gemma 4 the text tokenizer and audio encoder, and Google proposes it as the retrieval half of a device-side RAG pipeline. To pick the Gemma 4 variant based on your memory, check Gemma 4 on your Mac with Ollama: which variant to choose based on your RAM.

Google also mentions other ways to run it:

  • Libraries: Transformers, transformers.js (browser), MLX, vLLM, SGLang
  • Local runtimes: llama.cpp (official GGUF), LM Studio
  • On-device: LiteRT and MediaPipe

How much RAM does EmbeddingGemma 2 need?

Google reports about 191 MB of active RAM for text-only weights and about 567 MB for the full multimodal model. Both figures were measured with quantization on a Pixel 11 Pro, so they’re the best case.

In full precision or bfloat16, on a laptop or server, plan for more. The Ollama download sizes from the previous section give a more realistic idea of what it takes on a desktop machine.

What errors break your embeddings without warning?

Three, and none throw an exception.

Don’t use float16. The model’s activations exceed the float16 range. In float16 it returns NaN or degraded vectors silently, instead of failing. Use bfloat16 where your hardware supports it, and if not, float32, which covers most CPUs.

Re-normalize after truncating. EmbeddingGemma 2 uses Matryoshka Representation Learning (matryoshka embeddings), so you can trim vectors from 768 to 512, 256, or 128 dimensions. But a truncated vector no longer has unit length. Let the library do both steps:

emb = model.encode(texts, truncate_dim=256, normalize_embeddings=True)

Also, queries and documents have to use the same dimension.

Don’t truncate code too much. Google’s own table shows where the cost is:

Dimensions Storage MTEB Code
768 1× 78.68
512 1.5× less 77.24
256 3× less 76.18
128 6× less 71.41

256 is the sweet spot: a third of the storage for about 2.5 points of precision on code. With 128 dimensions, code drops nearly to version 1 levels, and Google warns that multimodal quality degrades considerably.

Which vector database should I use to store embeddings?

Any vector database that accepts vectors of the dimension you choose: EmbeddingGemma 2 only generates the vectors, it doesn’t store them. Among its launch partners, Google names Qdrant for storing them.

What does depend on the model is how much your index takes up. Choosing the dimension is the most important storage decision: storing at 256 dimensions instead of 768 reduces index size to a third, and per the table above barely costs precision. Decide on dimension before indexing: if you store truncated vectors and later want more dimensions, you’ll have to regenerate all embeddings from your corpus.

Is EmbeddingGemma 2 free and where do you download it?

Yes: the weights are open under Apache 2.0, a license that allows commercial use, so there’s no cost per token. What you pay for is compute on your own machine.

You can download it from:

  • Hugging Face: google/embeddinggemma-2
  • Kaggle
  • Ollama

At the time of publishing this note, Google listed availability in Model Garden as “coming soon”.

The model comes with documentation for inference, multimodal embeddings, and fine-tuning, plus a RAG quickstart notebook that combines it with Gemma 4.


What about you? What embedding model do you use today for RAG on your code?

Where do you generate your embeddings today?
  • With an API in the cloud
  • With a local model (Ollama, llama.cpp…)
  • It depends on the project, I use both
  • I prefer another option (tell us which)
0 votantes

Comment below or, if this article reached you by email, reply directly to the email: your reply gets published here.

Related: Improving RAG in Spanish: testing Mistral’s new embeddings