Google has introduced Gemma 4, the latest generation of its family of open-source models, in four sizes designed to cover everything from running on smartphones to workstation-level deployments. The models are built with the same research and technology that underpins Gemini 3, Google’s proprietary cutting-edge model, and are released under the Apache 2.0, a more permissive term than previous Gemma generations, a change that Hugging Face co-founder Clément Delanguedescribed as “a huge milestone.”
See also: Google Cast support is rolling out to Samsung TVs

Google DeepMind CEO Demis Hassabiscalled the new models “the best open models in the world for their respective sizes.” The four variants are the Effective 2B (E2B) and Effective 4B (E4B), designed to run on devices like phones, the Raspberry Pi, and the Jetson Nano, developed in collaboration with the Pixel team, Qualcomm, and MediaTek.
The 26B Mixture-of-Experts (MoE) and 31B Dense models are aimed at offline use on developer hardware and consumer GPUs. The 31B Dense model is currently ranked third among all open models in the Arena AI text leaderboard, while the 26B MoE is in sixth place.
Google claims that both larger models outperform models up to 20 times larger in this benchmark. The 31B unquantized weights fit on a single Nvidia H100 80GB GPU, while the quantized versions run on consumer hardware. All four models are multimodal, natively process video and images, and have been trained on over 140 languages.
See also: Google Home update: Better command understanding from Gemini

The E2B and E4B models also support native audio input for speech recognition. Content windows are 128K tokens for the edge models and 256K for the two larger variants. In terms of features, Google highlights improvements in multi-step logic, native function calls, and structured JSON output for agentic workflows and offline code generation. In terms of performance, the Android Developers Blog notes that the E2B model runs three times faster than the E4B, while the edge family as a whole is up to four times faster than previous Gemma versions and uses up to 60% less battery.
The E2B and E4B models also form the basis for the Gemini Nano 4, Google’s next-generation Android model, which will arrive on consumer devices later this year. Gemma has amassed more than 400 million downloads and over 100,000 community-created variants since its initial release, something Google points to as evidence of developer adoption at scale.
See also: Google Drive: AI ransomware detection for all subscribers

Gemma 4 is available immediately on Hugging Face, Kaggle , and Ollama, with the 31B and 26B models accessible via Google AI Studio and the edge models via the AI Edge Gallery. The decision to license to Apache 2.0 is the most significant trademark in the release: it removes restrictions that prevented some enterprise and commercial deployments under the previous Gemma terms, opening the ecosystem to a wider range of production use cases.
