Google Accelerates Gemini Nano Models on Pixel with Frozen Multi-Token Prediction
Google has introduced a method to retrofit Multi-Token Prediction (MTP) onto frozen production models like Gemini Nano v3, enabling faster on-device inference on Pixel 9 and 10 devices. The MTP head attaches to the main model's final layers, leveraging its hidden states and KV cache to generate multiple tokens per inference pass without separate drafting models, achieving speedups of 50% or more and reducing memory consumption by 130MB per instance.
Google/DeepMind
Google DeepMind
