EmbeddingGemma 2 is Google DeepMind's free, open model that turns text, code, images, video and audio into numbers you can search, and it's small enough to run on a laptop or phone. Google released it on October 6, 2026, under the Apache 2.0 license, which lets you use it in commercial projects.
Ever wanted a search box that gets what you mean, works offline and never sends your files to a server? This is the kind of model that makes it possible. Below you'll see what changed from the first EmbeddingGemma and when to pick it over Google's cloud model, Gemini Embedding 2. Then you'll run a short local search yourself.
First, what is an embedding?
Think about how you'd sort a pile of photos and notes by topic. You'd put the beach photos near the note that says "vacation," even though the words never match. An embedding model does the same thing with math.
It reads a piece of content and returns a list of numbers, called a vector, that captures its meaning. Content with similar meaning gets similar numbers. So when you search, the model turns your question into numbers too and finds the closest matches. That's semantic search: matching on meaning, not exact words.
This is also the first half of retrieval augmented generation (RAG). RAG means you look up the right documents first, then hand them to a chat model so it answers from your own material. Our guide to what RAG is walks through the whole idea.
What's new in EmbeddingGemma 2
The first EmbeddingGemma handled text only. Version 2 puts text, code, images, video and audio into one shared space. So a typed question can find a photo, a video moment or a voice recording. Here's how the two compare, based on Google's EmbeddingGemma docs and the EmbeddingGemma 2 model card.
| EmbeddingGemma 1 | EmbeddingGemma 2 | |
|---|---|---|
| What it reads | Text | Text, code, images, video and audio |
| Size | 308M parameters | 740M total, or 270M if you load text only |
| Built on | Gemma 3 | Gemma 4 |
| Context window | 2K tokens | 8,192 tokens |
| Output | 768 numbers per item | 768, or shortened to 512, 256 or 128 |
| License | Open weights | Apache 2.0 |
Three changes matter most in practice.
It's modular. You only load the parts you need. Text and code alone use 270M parameters. Add vision and it's 440M. Add audio and it's 570M. Load everything and it's 740M. Google's developer guide says all four setups share one vector space, so a text-only query can match documents you embedded with the full model.
It's better at code. Google reports a jump on the MTEB Code benchmark from 68.76 to 78.68, per its launch post. That makes it a fit for searching your own codebase.
It can shrink its output. A training method called Matryoshka Representation Learning puts the most important information first, so you can keep only the first 512, 256 or 128 numbers. Google's example: a million vectors take about 1.5 GB at full size and about 250 MB at 128, a 6x saving. The tradeoff is quality. Google says 256 keeps most of the quality on text and code and about 95% on image, video and speech search, while 128 drops media search to around 75%.
Should you use EmbeddingGemma 2 or Gemini Embedding 2?
Google now has two multimodal embedding models, and they solve different problems.
| EmbeddingGemma 2 | Gemini Embedding 2 | |
|---|---|---|
| Where it runs | Your own laptop, phone or server | Google's cloud, through the Gemini API |
| Internet needed | No | Yes |
| Your data leaves the device | No | Yes, it's sent to the API |
| Default output | 768 numbers | 3,072 numbers, which you can shorten |
| Weights | Open, Apache 2.0 | Closed |
Pick EmbeddingGemma 2 when privacy matters, when the app has to work offline, or when you want no per-call API bill. Think a notes app that searches your files on the phone, or a tool that indexes a private codebase.
Pick Gemini Embedding 2 when you'd rather not run models yourself and you're already building on the Gemini API. Its details come from Google's embeddings documentation.
One rule applies either way. Vectors from different models don't mix, so embed your documents and your queries with the same model.
Try it: a tiny local search in Python
This uses the sentence-transformers library, which is how Google's own Sentence Transformers guide runs the model. It loads the text-only setup to save memory. The first run downloads the model from Hugging Face, and after that it works offline.
Step 1. Install the libraries.
pip install -U sentence-transformers transformersStep 2. Load the model with text only and embed a few notes.
from sentence_transformers import SentenceTransformer
model = SentenceTransformer(
"google/embeddinggemma-2",
config_kwargs={"vision_config": None, "audio_config": None},
)
notes = [
"Invoice from the design agency is due on Friday.",
"Team offsite moved to the lake house in March.",
"Reset the staging database before the demo.",
]
note_vectors = model.encode(notes, prompt_name="Retrieval-document")Step 3. Ask a question and rank the notes.
question = "When do we have to pay the designers?"
question_vector = model.encode(question, prompt_name="Retrieval-query")
scores = model.similarity(question_vector, note_vectors)[0]
for note, score in sorted(zip(notes, scores), key=lambda x: -x[1]):
print(round(float(score), 3), note)How to check it worked: the invoice note should print first, even though the question never says "invoice." That's semantic search doing its job.
A few details from Google's guide that save time. Use the Retrieval-query prompt for questions and Retrieval-document for the content you store, because the model was trained to treat them differently. Don't load it in float16, since Google says the model doesn't support it, so stick with the default, bfloat16 or float32. And if you need smaller vectors, add truncate_dim=256 when you load the model.
Where you'll see it on phones and Macs
You don't have to write code to see what it does. Google's AI Edge post describes two demos in the Google AI Edge Gallery app for Android and iOS. Instant Media Search finds photos as you type. Video Moments Finder jumps to the part of a local video that matches your words. On the Mac, Google's experimental Foresight app uses it with Gemma 4 to search your meeting notes and files locally. Google also says the model will come to Android through ML Kit in the coming weeks, so check the status before you plan around it.
What it won't do
- It doesn't answer questions. It finds relevant content. To write an answer, you pair it with a chat model, which is the full RAG setup.
- Small vectors cost media quality. At 128 numbers, Google's own figures show image, video and speech search dropping to around 75% quality. Test on your own data first.
- It's brand new. Tooling around it, like ML Kit support, is still rolling out.
Go deeper
Embeddings are also how many AI agents remember things between sessions. Our guide on how AI agent memory works shows where search fits. And for more tools worth knowing, browse our AI tools roundup.
Want to build projects like this with a guided path instead of piecing it together alone? AI Flex lets you learn AI at your own pace, with an AI tutor, mentors and no fixed schedule. Explore how it works and see if it fits.
