
Google claims EmbeddingGemma 2 outperforms rival embedding models twice its size
Quick Answer
Google's EmbeddingGemma 2, with 740 million parameters, outperforms larger models in multimodal embedding tasks, scoring 78.68 on the Massive Text Embedding Benchmark.
Quick Take
It runs locally, requires minimal RAM, and supports offline applications, making it ideal for developers seeking efficient solutions.
Key Points
- EmbeddingGemma 2 has 740 million parameters, making it highly compact.
- It scores 78.68 on the Massive Text Embedding Benchmark, a significant improvement.
- The model requires only 191 MB of RAM and reduces local storage needs by six times.
- It can run offline applications with small models like Gemma 4 without external data transfer.
- Weights and documentation are available on Hugging Face and Kaggle.
📖 Reader Mode
~1 min readGoogle released EmbeddingGemma 2, an open model that converts text, images, video, audio, and code into numerical vectors so similar content can be found and compared more easily. At 740 million parameters, Google says it's the most compact model of its kind and outperforms competing models up to twice its size on multimodal embedding benchmarks.

The model runs locally without an API key. Each query takes about 20 to 70 milliseconds via WebGPU in the browser. It needs only around 191 MB of RAM and cuts local vector database storage by up to six times. For text-only tasks, a 270-million-parameter version is enough.
Paired with small open models like Gemma 4, EmbeddingGemma 2 can run offline RAG apps without sending data to external servers. The weights are available on Hugging Face and Kaggle, along with a developer guide and documentation.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

