Most AI search projects start with a hard question: do we have to send our files to someone else’s cloud to make them searchable? EmbeddingGemma 2 offers a different answer. An embedding model turns each document, photo or recording into a list of numbers so that similar things can be found together. Google’s new model does that for five kinds of content and is small enough to run on a laptop, a server in your own rack or, Google says, a phone.
What Google released
Google says the first EmbeddingGemma, a text-only model, passed 20 million downloads. Version 2 adds code, images, video and audio, and Google says it is built from the same technology as its Gemini Embedding models. The design is modular: load only the text backbone for document search, then add the vision or audio encoder when you need them. The Hugging Face card lists effective sizes of 270M for text only, 440M for text and image, 570M for text and audio, and 740M for the full model.
“Generating embeddings locally helps ensure data privacy, reduces pipeline latency, and empowers developers to build cross-modal search and retrieval that works entirely offline.”Google, EmbeddingGemma 2: an open, lightweight multimodal embedding model, October 6, 2026
Google’s performance claims include a 9.92-point gain on MTEB Code, from 68.76 to 78.68, and what it calls leading scores among sub-1B multimodal embedders. These are Google’s benchmarks, not a test on Alberta data. For storage, the model uses Matryoshka Representation Learning, so vectors can be cut from 768 to 512, 256 or 128 dimensions, for what Google calls up to 6x storage reduction. The model card warns that 128 dimensions degrades multimodal quality substantially and should be validated first.
Why it matters for private search
Search needs two parts: an embedding model that turns content into vectors, and an index that stores them. If both run on hardware you control, documents and media never have to leave the building to be searched. Google reports that a quantized EmbeddingGemma 2 needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro, which suggests modest hardware is enough for a pilot. Google lists support in common self-hosting tools, including sentence-transformers, vLLM, llama.cpp, Ollama and the Qdrant vector database.
| Configuration | Effective size | Example use |
|---|---|---|
| Text only | 270M | Procedures, reports, standards, code |
| Text and image | 440M | Inspection photos, drawings, scanned forms |
| Text and audio | 570M | Field voice notes, recorded meetings |
| Full multimodal | 740M | Mixed photo, video and audio archives |
Where Alberta teams could use it (our advice)
What follows is our advice, not Google’s. Industrial operators collect large amounts of media that nobody can search: drone and inspection photos, pipeline and tank videos, and voice notes from the field. A small, self-hosted search layer could let an engineer type “corrosion near a weld at a valve” and get matching photos and video moments back, without sending any of it to an outside service. The same setup works for private document search over procedures, contracts and incident reports.
Start small. Pick one archive, such as last year’s inspection photos for one site. Index it on a server you control, and test 30 real questions from your inspectors before expanding. Follow the model card’s precision advice: run in bfloat16 or float32, not float16. Keep your access controls in the search layer so people only find files they’re already allowed to open. The same discipline helps with AI search optimization on your public site: content that is well structured and clearly labelled is easier for any search system to find. Our industrial AI Alberta page covers field and plant use cases, and private AI security covers keeping that data in-house.
