Google DeepMind

Latest AI news, models and releases from Google DeepMind. ['3.5 Flash', '3.6 Flash', 'Aletheia', 'AlphaEarth Foundations', 'AlphaEvolve', 'AlphaFold', 'AlphaFold 3', 'AlphaGenome', 'AlphaGo', 'AlphaProof', 'AlphaZero', 'Antigravity 2.0', 'Auto frame', 'Co-Scientist', 'DiffusionGemma', 'ERA', 'Gemini', 'Gemini 2.5 computer use', 'Gemini3', 'Gemini 3.1 Flash-Lite', 'Gemini 3.5', 'Gemini 3.5 Flash', 'Gemini 3.5 Flash Cyber', 'Gemini 3.5 Flash-Lite', 'Gemini 3.5 Pro', 'Gemini 3.6 Flash', 'Gemini 4', 'Gemini Deep Research', 'Gemini Deep Think', 'Gemini for Education', 'Gemini for Science', 'Gemini Omni', 'Gemini Omni Flash', 'Gemini Spark', 'Gemma', 'Gemma 4', 'Gemma 4 12B', 'Gemma 4 26B', 'Gemma 4 26B MoE', 'gemma4:31b']

Google Accelerates Gemini Nano Models on Pixel with Frozen Multi-Token Prediction

Google has introduced a method to retrofit Multi-Token Prediction (MTP) onto frozen production models like Gemini Nano v3, enabling faster on-device inference on Pixel 9 and 10 devices. The MTP head attaches to the main model's final layers, leveraging its hidden states and KV cache to generate multiple tokens per inference pass without separate drafting models, achieving speedups of 50% or more and reducing memory consumption by 130MB per instance.

Google/DeepMindGoogle/DeepMind Google DeepMindGoogle DeepMind
Google Research27.07 · 05:05
Fresh news