Google DeepMind introduces Gemma 4 12B: a unified multimodal model without encoders
Google DeepMind has unveiled Gemma 4 12B, a mid-sized multimodal AI model that processes vision and audio directly without separate encoders. Designed to run on laptops with 16GB RAM, it offers advanced reasoning and agentic capabilities. The model is released under Apache 2.0 license and supports developer tools like Hugging Face, Ollama, and Google Cloud.
Google/DeepMind
Google DeepMind
