ModelsOpen Source 🇺🇸 27.07.2026 12:03

Google DeepMind introduces Gemma 4 12B: a unified multimodal model without encoders

Google/DeepMindGoogle/DeepMind Google DeepMindGoogle DeepMind
Google DeepMind has unveiled Gemma 4 12B, a mid-sized multimodal AI model that processes vision and audio directly without separate encoders. Designed to run on laptops with 16GB RAM, it offers advanced reasoning and agentic capabilities. The model is released under Apache 2.0 license and supports developer tools like Hugging Face, Ollama, and Google Cloud.
Google DeepMind announced Gemma 4 12B, a unified multimodal model without dedicated encoders for vision and audio inputs, allowing direct processing by the LLM backbone. It is designed for local execution on consumer laptops with 16GB of VRAM or unified memory, delivering performance close to the larger 26B MoE model but with a smaller memory footprint. Key features include native audio inputs, Multi-Token Prediction drafters for reduced latency, and an Apache 2.0 license. The model is available for download on Hugging Face and Kaggle, and supports integration with tools like LM Studio, Ollama, and Google Cloud. Gemma 4 models have surpassed 150 million downloads.
Сокращения
MoE = Mixture of Experts — смесь экспертов
LLM = Large Language Model — большая языковая модель
VRAM = Video Random Access Memory — видеопамять
MTP = Multi-Token Prediction — многотокенное предсказание
Source: Google DeepMind — original
Our earlier posts on this topic ↓
Fresh news