EmbeddingGemma 2: Run Google's Multimodal Embedding Model Locally (Ollama, Python, Kali) and on AWS

Logeshwaran
—

EmbeddingGemma 2 is Google DeepMind's new open embedding model, released on October 6, 2026, and it does something unusual for a model this small: it puts text, code, images, video and audio into one shared search space. Type "a cracked phone with a green line," and it can find the matching photo, the matching voice memo and the matching repair note, all with the same 768-number vectors. It has 740 million parameters, ships under Apache 2.0 with no sign-up gate, and runs on a phone. The full model is a 1.3 GB download in Ollama; the text-only version is 378 MB, smaller than a long 4K video. And one warning before you start: run it in float16 and it quietly returns broken vectors, with no error at all.

Jake runs a phone repair shop, and after last night's adventure with Mistral's trillion-parameter "le Chonk," which turned out to weigh about two terabytes, he was ready for a smaller cat. His problem is a very ordinary one. He has six years of repair photos, a few hundred voice memos he records while he works ("Galaxy, green line, customer says it fell in the sink"), and a spreadsheet of notes, and finding anything in them takes longer than the repair. When Ethan sent him the EmbeddingGemma 2 announcement, Jake asked one question: "Does it fit on my laptop this time?" Ethan grinned. "It fits on your phone. Your laptop will barely notice it's there." This page is what they did that evening: what EmbeddingGemma 2 is in plain English, how to run it with Ollama, Python, LM Studio or llama.cpp on Windows, Mac or Kali, how Jake built a private photo-and-voice-memo search, the silent mistakes that ruin results, and how to use it, or AWS's own alternative, in the cloud.

⚡ Quick Answer

• What is EmbeddingGemma 2? → Google's open embedding model for text, code, images, video and audio in one 768-dimension space. 740M parameters, 8,192-token context, Apache 2.0. What it is.

• Fastest way to run it → ollama pull embeddinggemma-2 (1.3 GB), or embeddinggemma-2:270m (378 MB) for text and code only. Ollama steps.

• For images, video and audio → use Python with sentence-transformers 6.1 or later and the model google/embeddinggemma-2. Working code.

• The silent trap → never run it in float16; use bfloat16 or float32. And re-normalize vectors after you shorten them. All the silent mistakes.

• On AWS? → Not on Bedrock. Self-host it on Lambda, EC2 or SageMaker AI, or use Amazon Nova Multimodal Embeddings, Bedrock's own multimodal embedder. AWS options and prices.

You do not need a GPU. The text-only setup runs comfortably on an ordinary laptop processor, and the full model is about the size of a phone game.

New to running AI on your own machine? Our plain-English series on running AI locally explains models, parameters and quantization in about ten minutes. You do not need it for this page, though. EmbeddingGemma 2 is the gentlest possible first local model: small, free, quick to download and genuinely useful on day one.

🧭 NEW HERE? READ THESE FIRST

New to local AI or AWS? These five pages make the rest of this one easy:

📌 Bookmark this; the setup commands and the "silent mistakes" list are the parts you will come back for.

What EmbeddingGemma 2 is, in plain English

An embedding model does not chat. It turns things into lists of numbers, called vectors, so that things with similar meaning get similar numbers. A photo of a cracked Galaxy screen and the sentence "Samsung with a broken display" end up close together; a photo of a birthday cake ends up far away. Once everything is a vector, search becomes simple arithmetic: turn the question into a vector and find the nearest ones. That is how "semantic search" works, and it is the retrieval half of every RAG system, the setups where a chat model answers questions using your own documents.

What makes EmbeddingGemma 2 unusual is the "everything" part. Until now, a small open model usually handled one kind of input: text models for text, image models for images, and you glued them together and hoped. EmbeddingGemma 2 maps text, code, images, video and audio, and even mixtures of them, into one shared space. A text query can find a video clip. A voice memo can find a photo. A product listing with words, two photos and a short video becomes one single vector.

The facts, from Google's model card:

  • Size: 740 million parameters in total, built in modules: a 270M text model (a 130M transformer plus a 140M embedder), a 170M vision encoder and a 300M audio encoder. You load only the parts you need.
  • Output: 768-dimension vectors, which you can shorten to 512, 256 or 128 dimensions to save storage.
  • Context: 8,192 tokens per input, four times EmbeddingGemma 1. That is about 5.5 minutes of audio, 29 images or 58 video frames in one input.
  • Languages: more than 100, trained on data in over 140, plus code.
  • Built on: Gemma 4, sharing its text tokenizer and audio encoder, which matters when you run the two together.
  • License: Apache 2.0, free for commercial use. On Hugging Face the model is not gated, so you can download it without signing in or accepting extra terms.
  • Training data cutoff: January 2025.

Google says EmbeddingGemma 1 was downloaded more than 20 million times, mostly by people building private, on-device search. The new version keeps that model's text quality and adds everything else.

EmbeddingGemma 2 vs EmbeddingGemma 1

EmbeddingGemma 2EmbeddingGemma 1
ReleasedOctober 6, 20262025
Hugging Face namegoogle/embeddinggemma-2google/embeddinggemma-300m
Parameters740M (270M for text only)About 300M
InputsText, code, images, video, audioText and code
Context8,192 tokensAbout 2,000 tokens
MTEB multilingual v261.3661.15
MTEB code v178.6868.76
License on Hugging FaceApache 2.0, not gatedGemma license, gated (accept terms first)

Two things to take from that table. For plain text, version 2 is about the same quality as version 1, so do not expect your text search to transform overnight. For code, it is nearly ten points better, a jump of about 14%, which makes it a real option for searching your own codebase. And if you already have an index built with version 1, the vectors are not interchangeable. Switching models means embedding your collection again.

How good is it outside text?

Google's model card reports, at full 768 dimensions: 64.64 on the image benchmark MIEB lite, 57.28 on MMEB v2 image tasks, 67.84 on visual documents such as PDFs and slides, 50.67 on video, 69.54 on the MSEB audio retrieval benchmark and 49.39 on MAEB. On their own those numbers mean little to most of us. What they add up to, in Google's words, is the best quality for its size among multimodal embedders under 1 billion parameters, and better than some specialist models more than twice as big. The honest test, as always, is twenty of your own searches. Jake's was "green line," "water damage" and "the iPad with the dog bite," and it found all three.

EmbeddingGemma 2 sizes: how much memory and disk it needs

This is the part where Jake relaxed. Every number below is small, and the modular design means you only pay for what you use.

What you loadParametersOllama tag and downloadUse it for
Text and code only270Membeddinggemma-2:270m, 378 MBDocument search, RAG, code search
Text plus images and video440Membeddinggemma-2:440m, 714 MBPhoto libraries, scanned documents, product catalogs
Text plus audio570Membeddinggemma-2:570m, 990 MBVoice memos, podcasts, call recordings
Everything740Membeddinggemma-2:740m (also latest), 1.3 GBMixed media search

Other formats, for reference. The full-precision weights on Hugging Face are a single 1.49 GB file. The GGUF builds for llama.cpp come in two parts: the text model (558 MB in BF16, 310 MB in Q8_0) and a separate file for the vision and audio encoders (982 MB in BF16, 555 MB in Q8_0). On a phone, Google measured about 191 MB of active memory for the quantized text-only model on a Pixel 11 Pro, and about 567 MB for the full multimodal model.

Translated into hardware: any laptop from the last several years will run the text-only model, on the processor alone. 8 GB of RAM is enough for the text setup, 16 GB is comfortable for the full model alongside your browser, and a GPU only matters when you are embedding tens of thousands of files and want it done before lunch. If you are also planning to run a chat model next to it, our honest laptop guide for local AI has the tiers.

The vectors themselves need space too, and this is where shortening them pays off. Google's developer guide puts it simply: a million 768-dimension vectors in bfloat16 take about 1.5 GB, while the same million at 128 dimensions take about 250 MB. That is a sixfold saving, and for a phone app or a small server it is the difference between fitting and not fitting.

How to run EmbeddingGemma 2 with Ollama (Windows, Mac, Linux, Kali)

Ollama is the quickest way in. If you do not have it yet, install it from ollama.com on Windows or Mac; on Linux or Kali, the one-line installer from the same site works. If you are choosing between local AI apps, our comparison of Ollama, LM Studio, Jan and the rest explains the trade-offs.

  1. Update Ollama, then pull the model. This model needs Ollama 0.36.0 or later. For everything: ollama pull embeddinggemma-2. For text and code only, which is smaller and faster: ollama pull embeddinggemma-2:270m.
  2. Make sure Ollama is running. The desktop app starts the server for you; on Linux you can run ollama serve. If you see "could not connect," our fix for Ollama connection errors covers every cause.
  3. Ask for an embedding with the command below. You get back a list of 768 numbers, which is the vector.
  4. Store the vectors in whatever you like: a simple file, SQLite, or a vector database. Then compare a query vector against them with cosine similarity.
curl http://localhost:11434/api/embed -d '{
  "model": "embeddinggemma-2",
  "input": "task: search result | query: phone screen with a green line"
}'

Notice the strange text at the start of that input. That is a task prefix, and it matters. EmbeddingGemma 2 was trained with short instructions in front of text, and it produces better vectors when you include them. Ollama's examples send plain text, and the model applies no prefix by default, so you type them yourself:

  • Search queries: task: search result | query: {your question}
  • Documents being searched: title: {title} | text: {content}, or title: none | text: {content} when there is no title.
  • Code search queries: task: code retrieval | query: {your question}, with code files stored as title: {filename} | text: {code}.
  • Classification, clustering and similarity: the same prefix on every input, for example task: clustering | query: {content}.

The same call from Python with the official ollama package:

import ollama

docs = [
    "title: none | text: Galaxy S24, green vertical line after a drop, display replaced.",
    "title: none | text: iPad 10, charging port full of lint, cleaned, no parts needed.",
]
query = "task: search result | query: phone screen with a green line"

doc_vecs = ollama.embed(model="embeddinggemma-2:270m", input=docs).embeddings
query_vec = ollama.embed(model="embeddinggemma-2:270m", input=query).embeddings[0]

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    norm = (sum(x * x for x in a) ** 0.5) * (sum(y * y for y in b) ** 0.5)
    return dot / norm

for doc, vec in zip(docs, doc_vecs):
    print(round(cosine(query_vec, vec), 3), doc)

The Galaxy note is the one that should score highest, and that is the whole idea in a dozen lines. One small thing that will confuse people this week: Ollama's model page lists every EmbeddingGemma 2 tag with a "256K context window." The model itself reads 8,192 tokens per input, which is what Google's model card says. Split long documents into chunks of a few hundred to a couple of thousand words, and you will never notice the difference.

How to use EmbeddingGemma 2 in Python (text, images, video and audio)

For anything beyond text, use Python with Hugging Face's sentence-transformers library. It handles the task prefixes for you, loads only the encoders you ask for, and supports every modality. You need sentence-transformers 6.1.0 or later.

Install it (Windows, Mac or Kali)

Create a virtual environment first, so this install cannot break anything else on your system:

python3 -m venv ~/eg2
source ~/eg2/bin/activate
pip install -U "sentence-transformers[image,audio,video]" transformers

On Windows, the activate line is eg2\Scripts\activate instead. Two details save Kali users an evening. First, keep the quotes around the package name: Kali's default shell is zsh, and without quotes zsh reads the square brackets as a file pattern and answers "no matches found." Second, if you skip the virtual environment, Kali refuses with "externally-managed-environment." That is Kali protecting itself, not a bug, and our guide to the externally-managed-environment error explains the right fix.

Text and code search

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("google/embeddinggemma-2")

query = "What causes the northern lights?"
document = "The northern lights are caused by charged particles from the sun."

query_emb = model.encode(query, prompt_name="SearchQuery")
doc_emb = model.encode(document, prompt_name="Document")
print(model.similarity(query_emb, doc_emb))

prompt_name adds the right task prefix for you. The useful names are SearchQuery and Document for search, QuestionAnswering, FactChecking, CodeRetrieval, and Classification, Clustering and SentenceSimilarity for the symmetric tasks. Document uses "title: none"; if your documents have real titles, format them yourself as title: {title} | text: {content}.

Load only what you need

The first line above loads all 740M parameters. If you only work with text, tell it to skip the vision and audio encoders. They are never loaded, so you save the memory entirely:

MODEL_ID = "google/embeddinggemma-2"

# Text and code only (270M)
text_model = SentenceTransformer(MODEL_ID, config_kwargs={"vision_config": None, "audio_config": None})

# Text, images and video (440M)
image_model = SentenceTransformer(MODEL_ID, config_kwargs={"audio_config": None})

# Text and audio (570M)
audio_model = SentenceTransformer(MODEL_ID, config_kwargs={"vision_config": None})

Here is the clever part. All four setups share one vector space. A query embedded with the 270M text-only setup can be compared directly with documents embedded by the full model. Start text-only today, add photos next month, and the vectors you already have stay valid.

Images, audio and video

Pass media as a dictionary keyed by its type, without any prompt. Prefixes are for text only:

image_emb = model.encode({"image": "sunset_beach.jpg"})
audio_emb = model.encode({"audio": "ocean_waves.wav"})
query_emb = model.encode("ocean waves at sunset", prompt_name="SearchQuery")

print(model.similarity(query_emb, image_emb))
print(model.similarity(query_emb, audio_emb))

You can also mix them into one input. Mark where each item goes with <|image|>, <|video|> or <|audio|>, and you get a single vector for the whole thing:

listing_emb = model.encode({
    "text": "Waterproof trail shoe. <|image|> Grip test on wet rock: <|video|>",
    "image": "trail_shoe.jpg",
    "video": "grip_test.mp4",
})
query_emb = model.encode("waterproof trail shoes", prompt_name="SearchQuery")
print(model.similarity(query_emb, listing_emb))

Media can be file paths (MP4 for video), URLs for images and audio, or in-memory images, arrays and tensors. Audio should be mono at 16 kHz. Video is sampled at one frame per second by default. Everything in one input shares the 8,192-token budget: an image costs 280 tokens by default, a video frame 140, and a second of audio 25. A lower vision budget fits up to about 114 images or frames per input, at some cost to detail.

Shorter vectors with Matryoshka

EmbeddingGemma 2 is trained so the first part of each vector carries most of the meaning, like Russian nesting dolls, hence the name Matryoshka. Ask for a shorter vector and let the library re-normalize it:

query_emb = model.encode(
    query,
    prompt_name="SearchQuery",
    truncate_dim=256,          # or 512, 128
    normalize_embeddings=True,
)

How much quality you give up, from Google's own table:

DimensionsStorageText (MTEB multilingual)Code (MTEB code)Multimodal (MMEB overall)
768 (full)1×61.3678.6859.01
5121.5× smaller61.1777.2458.38
2563× smaller60.4176.1856.24
1286× smaller57.8971.4145.65

The rule of thumb from Google's guide: 768 or 512 for multimodal and visual documents, 256 when storage is tight, and 128 only for large text-only indexes or as a first rough shortlist before a better re-ranking step. At 128, text keeps around 90% of its quality but images, video and speech drop to around 75%.

What you can build with EmbeddingGemma 2

"Embeddings" sounds abstract until you see the list of everyday jobs it does. Google's model card names five families of use, and every one of them works offline with this model:

  • Search and retrieval: document search, RAG over a company knowledge base, finding code from a plain-English question, or searching an archive of recordings by speaking a query.
  • Classification: sorting inputs into labels you choose, such as positive or negative reviews, content moderation, recognizing sounds in audio, or tagging photos by category.
  • Clustering: grouping things that mean the same, such as customer feedback by theme, a messy document folder by topic, or near-duplicate files across formats, like the same product shown in a photo and described in a listing.
  • Similarity: recommendations ("more like this"), duplicate detection, paraphrase spotting, and matching the same content across languages.
  • Fact verification: given a claim, finding the documents that support or contradict it, the first step of any fact-checking or source-citing tool.

Google's launch demos show the multimodal side best: searching a media library with either a sentence or a photo, finding the exact moment in a video from a text or spoken query, and pairing EmbeddingGemma 2 with Gemma 4 so a device can answer questions about the files stored on it. There is even a drawing game, the Embedding Draw Challenge, in which the model guesses what you are drawing. It is a surprisingly good way to understand what "a shared space for images and text" means in practice.

For a small business, the realistic first projects are boring in the best way: a search box over years of invoices and photos, an inbox that tags incoming messages, or a help page that finds the right answer even when the customer uses different words. None of them needs a cloud account, and all of them run on the laptop you already own.

A real example: Jake's private photo and voice-memo search

Here is the little tool Ethan built for Jake that evening, in about forty lines. It indexes a folder of repair photos and a folder of voice memos, then answers plain-English questions across both. Nothing leaves the laptop, which matters when the photos show customers' phones, and sometimes customers' faces.

Step one is the voice memos. Phone recordings are usually stereo and 44.1 or 48 kHz, and the model wants mono at 16 kHz, so convert them once with ffmpeg:

mkdir -p memos16k
for f in memos/*.m4a; do ffmpeg -loglevel error -i "$f" -ac 1 -ar 16000 "memos16k/$(basename "${f%.*}").wav"; done

Step two is the index and the search:

from pathlib import Path
import numpy as np
import torch
from sentence_transformers import SentenceTransformer

# bfloat16 on GPUs that support it, float32 everywhere else. Never float16.
dtype = torch.bfloat16 if torch.cuda.is_available() and torch.cuda.is_bf16_supported() else torch.float32
model = SentenceTransformer("google/embeddinggemma-2", model_kwargs={"torch_dtype": dtype})

photos = sorted(Path("photos").glob("*.jpg"))
memos = sorted(Path("memos16k").glob("*.wav"))
files = photos + memos

vectors = np.stack(
    [model.encode({"image": str(p)}, truncate_dim=512, normalize_embeddings=True) for p in photos]
    + [model.encode({"audio": str(m)}, truncate_dim=512, normalize_embeddings=True) for m in memos]
)
np.save("index.npy", vectors)

def search(question, top=5):
    q = model.encode(question, prompt_name="SearchQuery", truncate_dim=512, normalize_embeddings=True)
    scores = vectors @ q
    for i in np.argsort(-scores)[:top]:
        print(f"{scores[i]:.3f}  {files[i]}")

search("Samsung with a green line down the screen")
search("water damage, customer said it fell in the sink")

A few choices in there are deliberate. The vectors are normalized, so a plain dot product is the cosine similarity, and every vector, photos, memos and queries, uses the same 512 dimensions. The index is saved to a file, so it is built once and searched instantly afterwards. For a few thousand files, a NumPy array is all the "vector database" you need; when you reach hundreds of thousands, move to a real one.

Jake's first search, "green line," returned three photos and a voice memo from 2023 in which he could be heard saying "green line again, these screens hate water." He played it twice. "It found me," he said, slightly alarmed. Ethan pointed out that this was the point.

Add a chat model: on-device RAG with Gemma 4

Search finds the right files; a chat model can then read them and answer. EmbeddingGemma 2 was designed to pair with Gemma 4 on the same device. The two share a text tokenizer and an audio encoder, so running both costs less memory than two unrelated models would. The pattern is simple: embed the question, take the top few matches, and pass their text to Gemma 4 with the question. Our guides to running Gemma 4 on Windows and running Gemma 4 offline on Kali cover the chat half.

LM Studio, llama.cpp, MLX and the browser

Google worked with the usual local AI tools before launch, so you have choices beyond Ollama and Python:

  • LM Studio: search for EmbeddingGemma 2 in the model search, download it, and LM Studio's local server exposes it through the OpenAI-style /v1/embeddings endpoint. Handy if your app already speaks the OpenAI API. If LM Studio refuses to load a model, our LM Studio "failed to load model" guide walks through the causes.
  • llama.cpp: the official GGUF builds are in ggml-org/embeddinggemma-2-GGUF on Hugging Face. The text model is a single 310 MB file in Q8_0, and the vision and audio encoders come as a separate file you add only when you need them. Run it with llama.cpp's server in embedding mode.
  • MLX: Apple Silicon users get ready-made MLX versions from the mlx-community collection on Hugging Face.
  • vLLM and SGLang: both support it for serving at scale on GPUs.
  • In the browser: transformers.js and WebGPU run it client-side, and a WebGPU demo on Hugging Face shows it working with no server at all.
  • On phones: Google AI Edge, through MediaPipe for ready-made search and decision tasks or LiteRT for custom apps.
  • Fine-tuning: sentence-transformers and Unsloth both have guides for training it on your own data, which helps most for specialist vocabulary.

The silent mistakes: why your EmbeddingGemma 2 results look wrong

Embedding models rarely crash. They return numbers, the numbers look fine, and the search results are quietly worse. That makes these mistakes hard to spot, so here they all are in one place.

Running it in float16

This is the big one. EmbeddingGemma 2's internal values go beyond what float16 can hold. In float16 it returns NaN values or silently degraded embeddings instead of raising an error. Use bfloat16 on GPUs that support it, which also halves memory compared with float32, and float32 everywhere else, including most CPUs. Plenty of tutorials and scripts default to float16 for every model, so check yours.

Shortening vectors without re-normalizing

If you cut a 768-number vector down to 256 yourself by slicing it, the result is no longer unit length, and cosine scores drift. The ranking still looks plausible, which is exactly why it is dangerous. Use truncate_dim with normalize_embeddings=True, or normalize after slicing.

Comparing vectors of different lengths

A 768-dimension query cannot be scored against a 128-dimension index. Pick one length per index and use it for every document and every query.

Forgetting the task prefixes

Without prefixes the model still works, just less precisely. In sentence-transformers, use prompt_name. In Ollama, LM Studio or llama.cpp, type the prefixes into the text yourself. And never put a prefix on images, video or audio.

Mixing models in one index

Vectors from EmbeddingGemma 1, EmbeddingGemma 2, nomic-embed-text or a cloud model live in different spaces. Comparing them gives meaningless scores. When you switch models, embed the whole collection again. The only exception is EmbeddingGemma 2's own four setups, which share one space by design.

Trusting the "256K context" label

Ollama's page shows 256K for every tag; the model reads 8,192 tokens per input. Anything longer should be split into chunks, which gives better search results anyway.

Audio in the wrong format

The model expects mono audio at 16 kHz. Convert first with ffmpeg, as in Jake's example. Long recordings also hit the shared budget: about 327 seconds of audio fills a whole input, so split hour-long recordings into a few-minute pieces.

"No matches found" or "externally-managed-environment" on Kali

The first is zsh reading the square brackets in sentence-transformers[image,audio,video]; put the name in quotes. The second is Kali refusing a system-wide pip install; use a virtual environment.

Images or audio fail to load in Python

EmbeddingGemma 2's multimodal features need sentence-transformers 6.1.0 or later, plus the image, audio and video extras. Run pip install -U "sentence-transformers[image,audio,video]" transformers inside your virtual environment.

EmbeddingGemma 2 on AWS: Bedrock, Lambda, SageMaker and Nova Multimodal Embeddings

EmbeddingGemma 2 is not on Amazon Bedrock, and it cannot be brought there either: Bedrock's Custom Model Import feature does not support embedding models at all. On Google's own cloud, Google says availability in its Model Garden is "coming soon." So on AWS you have two honest paths: host EmbeddingGemma 2 yourself, or use AWS's own multimodal embedding model. New to Bedrock? Our plain-English guide to Amazon Bedrock explains the idea first.

Path 1: self-host EmbeddingGemma 2

Because the model is so small, self-hosting on AWS is unusually easy:

  • AWS Lambda for text-only embedding on demand. A Lambda function can have up to 10,240 MB of memory and run from a container image of up to 10 GB, so the 270M text setup fits with room to spare. Bake the model into the image so it is not downloaded on every cold start.
  • Amazon EC2 for steady or multimodal work. A small CPU instance handles text; a single modest GPU instance makes bulk image and video indexing fast.
  • Amazon SageMaker AI if you want a managed endpoint with autoscaling. Our explainer on Hugging Face models on AWS covers how Hugging Face models are deployed there.

A minimal Lambda container for text embeddings looks like this:

FROM public.ecr.aws/lambda/python:3.12
RUN pip install --no-cache-dir torch --index-url https://download.pytorch.org/whl/cpu \
 && pip install --no-cache-dir "sentence-transformers>=6.1.0" transformers
RUN python -c "from sentence_transformers import SentenceTransformer; SentenceTransformer('google/embeddinggemma-2', cache_folder='/opt/models')"
ENV HF_HUB_OFFLINE=1
COPY app.py ${LAMBDA_TASK_ROOT}
CMD ["app.handler"]
# app.py
import torch
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "google/embeddinggemma-2",
    cache_folder="/opt/models",
    config_kwargs={"vision_config": None, "audio_config": None},  # text only, 270M
    model_kwargs={"torch_dtype": torch.float32},                   # CPU: float32, never float16
)

def handler(event, context):
    vecs = model.encode(
        event["texts"],
        prompt_name=event.get("prompt_name", "Document"),
        truncate_dim=256,
        normalize_embeddings=True,
    )
    return {"embeddings": vecs.tolist()}

Give the function a few gigabytes of memory, which also gives it more CPU, and raise the timeout above the default 3 seconds; the first request after a cold start loads the model. Send queries with "prompt_name": "SearchQuery" and documents with the default.

Path 2: Amazon Nova Multimodal Embeddings on Bedrock

If you would rather not run anything, Bedrock has its own multimodal embedder: Amazon Nova Multimodal Embeddings, model ID amazon.nova-2-multimodal-embeddings-v1:0, launched October 28, 2025. Like EmbeddingGemma 2, it embeds text, images, video and audio. It is billed per use; on-demand prices in US East (N. Virginia):

InputNova Multimodal Embeddings priceExample jobExample cost
Text$0.135 per million tokens10 million tokens of documents$1.35
Standard image$0.00006 per image100,000 photos$6.00
Document image$0.0006 per image10,000 scanned pages$6.00
Video$0.0007 per second100 hours of video$252.00
Audio$0.00014 per second100 hours of recordings$50.40

Two things to know before you choose it. Its model card lists an earliest end-of-life date of October 28, 2026, three weeks from now, followed by a legacy period of at least six months, so check AWS's model lifecycle page before you build a large index on it. And Nova vectors and EmbeddingGemma 2 vectors are not compatible: pick one model per index, or plan to re-embed if you switch.

Which one to pick

For text and photos, the managed route is cheap. Jake's whole photo archive would cost him a few dollars to embed on Nova. The case for EmbeddingGemma 2 is not price but where the data goes and where the search runs: fully offline, on a laptop, on a phone, or inside your own AWS account with nothing sent to any model provider. For video, the arithmetic changes. A hundred hours costs about $252 on Nova each time you embed it, while EmbeddingGemma 2 on a GPU instance you already run costs only the instance time. If you are weighing Bedrock against hosting models yourself in general, our Bedrock vs SageMaker AI explainer lays out who manages what, and our Bedrock pricing guide explains the rest of the bill.

EmbeddingGemma 2 vs nomic-embed-text and other local embedders

If you already run a local embedding model, it is probably a text-only one such as nomic-embed-text, a popular default in Ollama tutorials, or EmbeddingGemma 1. Here is how to decide whether to switch:

  • Your data is only text: EmbeddingGemma 2's 270M setup is about as good as version 1 on general text and much better on code. If your current model works and you do not search code, there is no rush. If you search code, switch.
  • You have photos, scans, recordings or video: this is where EmbeddingGemma 2 has no small open rival. One model, one index, one search box, instead of separate tools stitched together.
  • You need commercial use without extra terms: Apache 2.0 is about as simple as licenses get, and the download is not gated.
  • You are building for phones or browsers: the quantized footprint of a few hundred megabytes and Google's own on-device tooling make it the obvious first choice.

Whatever you choose, remember the rule from the mistakes section: switching models means re-embedding everything. Plan an afternoon for it, not five minutes. Here is the order that avoids a broken search box in the middle:

  1. Save twenty real queries your users actually type, with the results you expect for each. This is your before-and-after test.
  2. Choose one vector length for the new index, for example 512 for mixed media or 256 for text, and write it down.
  3. Build the new index next to the old one. Re-embed every document with EmbeddingGemma 2, using the document prefix and the dimension you chose. Leave the old index running.
  4. Embed your saved queries with the new model, using the query prefix, and run them against the new index only.
  5. Compare the results with your expected answers. If the new index wins, or ties while adding photos and audio, you have your answer.
  6. Switch your app to the new model and index together, in one change, so queries and documents never come from different models. Then delete the old index.

EmbeddingGemma 2: frequently asked questions

What is EmbeddingGemma 2?

EmbeddingGemma 2 is an open embedding model from Google DeepMind, released on October 6, 2026. It turns text, code, images, video and audio into 768-dimension vectors in one shared space, for search, RAG, classification and clustering. It has 740M parameters and an Apache 2.0 license.

How do I run EmbeddingGemma 2 in Ollama?

Update Ollama to 0.36.0 or later, run ollama pull embeddinggemma-2, or embeddinggemma-2:270m for text only, then send text to the /api/embed endpoint on localhost port 11434. Add the task prefixes yourself, such as "task: search result | query:" before search queries.

Is EmbeddingGemma 2 free for commercial use?

Yes. It is released under the Apache 2.0 license, which allows commercial use, and the Hugging Face download is not gated. EmbeddingGemma 1 used the Gemma license and required accepting its terms first.

How much RAM does EmbeddingGemma 2 need?

Very little. The Ollama downloads are 378 MB for text only and 1.3 GB for the full model. An 8 GB laptop runs the text setup, and 16 GB is comfortable for the full model. On a Pixel 11 Pro, Google measured about 191 MB for quantized text-only and 567 MB for full multimodal.

Does EmbeddingGemma 2 need a GPU?

No. It runs on an ordinary laptop processor. A GPU only speeds up embedding very large collections, such as tens of thousands of photos or hours of video.

What is the EmbeddingGemma 2 context length?

8,192 tokens per input, shared across all media. That fits about 29 images, 58 video frames or 5.5 minutes of audio. Ollama's page shows 256K, but the model card's 8,192 is the real limit.

Can EmbeddingGemma 2 embed images and audio in Ollama?

Ollama's examples cover text through the /api/embed endpoint. For images, video, audio and mixed inputs, the documented route is Python with sentence-transformers 6.1.0 or later.

What are the EmbeddingGemma 2 task prefixes?

For search, use "task: search result | query: {query}" for queries and "title: {title} | text: {content}" for documents, with "title: none" if there is no title. Code search uses "task: code retrieval | query:". Clustering, classification and similarity use one shared prefix for all inputs.

Why does EmbeddingGemma 2 return NaN values?

You are almost certainly running it in float16, which cannot hold its activation range. Switch to bfloat16 on supported GPUs or float32 on CPUs. In sentence-transformers, set model_kwargs={"torch_dtype": torch.float32}.

Can I mix EmbeddingGemma 1 and EmbeddingGemma 2 vectors?

No. They are different models with different vector spaces, so scores between them are meaningless. Re-embed your collection when you switch. EmbeddingGemma 2's own text-only, image, audio and full setups do share one space.

What dimensions does EmbeddingGemma 2 support?

768 by default, and 512, 256 or 128 through Matryoshka truncation. 256 keeps most text and code quality at a third of the storage. 128 is best kept for text-only indexes, because image, video and audio quality drop to about 75%.

Is EmbeddingGemma 2 on AWS Bedrock?

No, and Bedrock's Custom Model Import does not support embedding models. Self-host it on Lambda, EC2 or SageMaker AI, or use Amazon Nova Multimodal Embeddings, Bedrock's own model for text, image, video and audio embeddings.

How do I install EmbeddingGemma 2 on Kali Linux?

Create a virtual environment with python3 -m venv, activate it, then run pip install -U "sentence-transformers[image,audio,video]" transformers with the quotes, which zsh needs. Or install Ollama and pull embeddinggemma-2.

Is EmbeddingGemma 2 better than nomic-embed-text?

For mixed media, yes by design, because nomic-embed-text handles text only. For plain text, test both on your own searches. EmbeddingGemma 2 is strongest on code, where it scores 78.68 on MTEB code.

Can I use EmbeddingGemma 2 with Gemma 4 for RAG?

Yes, and that pairing is what Google designed it for. The two share a text tokenizer and audio encoder, which lowers the combined memory. Embed and retrieve with EmbeddingGemma 2, then pass the top matches to Gemma 4 to answer.

Where can I download EmbeddingGemma 2?

The weights are on Hugging Face as google/embeddinggemma-2 and on Kaggle. Ollama has it as embeddinggemma-2, GGUF builds for llama.cpp are in ggml-org/embeddinggemma-2-GGUF, and Apple Silicon users can use the mlx-community versions.

By ten o'clock Jake's six years of photos and a few hundred voice memos were searchable from one box on his laptop, with nothing uploaded anywhere. He spent the next half hour typing things like "the phone that smelled of curry" and "the kid who dropped it in the fish tank," and laughing at what came back. The model doing all this was smaller than his holiday photo folder. Yesterday's model needed a server room; today's fits in a pocket. That is the other half of this AI year, and for most small businesses it may turn out to be the more useful half.

📌 If you keep one line from this page

One small model, one index, every kind of file, and never in float16.

Load only the encoders you need, keep one vector length per index, and re-embed everything when you change models.

Revision note. Written October 6, 2026, the day EmbeddingGemma 2 was released. If you have a folder of photos or recordings you gave up on searching years ago, tonight is a good night to try again.

Related