Gemma 4 E2B
Google DeepMind · United States · 2026
The smallest Gemma 4: 2.3 billion effective parameters, still with image and audio input, a 128K context and a built-in thinking mode.
Gemma 4 E2B is the entry point to the family and, together with E4B, was published on 2 March 2026, nine days ahead of the models that headlined the launch. It carries about 5.1 billion parameters on disk of which roughly 2.3 billion are effective in memory, the difference coming from Per-Layer Embeddings: each of its 35 decoder layers holds its own small embedding for every token, tables that are big but used only for lookups. The result is a model that behaves in memory like a two-billion-parameter model while retaining the vocabulary and the multimodal machinery of a much larger one. The feature list is remarkably complete for the size. It accepts text, images and audio and returns text, through a vision encoder of roughly 150 million parameters and an audio encoder of roughly 300 million, and it does automatic speech recognition and speech-to-translated-text. The context window is 128,000 tokens with 512-token sliding windows, and the vocabulary is the family standard 262,000. The configurable thinking mode is present here as in every other Gemma 4 - Google did not strip reasoning out of the small models. The scores are the ones to read carefully, because this is where a small model tells the truth about itself. E2B manages 60.0 percent on MMLU Pro, 37.5 on AIME 2026 without tools and 43.4 on GPQA Diamond; the last of those is essentially the level of chance-plus-knowledge and well below the 82.3 percent of the 26B. On document understanding it is the weakest of the family. Where it does convince is against history: on AIME it beats Gemma 3 27B, a model twelve times its effective size, by 37.5 to 20.8. Read on 25 August 2026 the instruction-tuned repository records 3.53 million downloads in thirty days.
▸Videos
No videos yet.
▸Reviews
No reviews yet. Be the first!