DiffusionGemma: 4x faster text generation
Google DeepMind has introduced DiffusionGemma, an experimental open 26B Mixture of Experts model that uses text diffusion to generate up to 4x faster than autoregressive LLMs on GPUs. It simultaneously processes 256-token blocks, targeting speed-critical local workflows like in-line editing, and is released under Apache 2.0 on Hugging Face.
Google/DeepMind
DeepMind
NVIDIA
