Google/DeepMind

Latest AI news, models and releases from Google/DeepMind. ['A2UI v0.9', 'AI Co-Scientist', 'AI Evaluator', 'AI Mode', 'AI Overviews', 'AlphaEvolve', 'AlphaFold', 'AlphaGenome', 'AlphaQubit', 'Antigravity', 'Auto frame', 'BERT', 'Chinchilla', 'Computational Discovery', 'Confidential GKE Nodes', 'Co-Scientist', 'Empirical Research Assistance', 'Era', 'ERA', 'Executive LLM', 'Farmscapes 2020', 'Flood Hub', 'frontier AI', 'Frozen v2', 'FunctionGemma', 'Gemini', 'Gemini 1.5 Flash', 'Gemini-1.5-pro', 'Gemini 1.5 Pro', 'Gemini 2.0 Flash', 'Gemini 2.5', 'Gemini 2.5 Flash', 'Gemini-2.5 Flash', 'Gemini-2.5-Flash', 'gemini-2.5-flash-image', 'Gemini 2.5 Pro', 'Gemini-2.5 Pro', 'Gemini-2.5-Pro', 'Gemini 3.1 Flash', 'gemini-3.1-flash-image']

Research 🇺🇸

LLM Architecture Innovations: KV Sharing, Per-Layer Embeddings, and Compressed Attention

Sebastian Raschka reviews recent open-weight LLM architecture advances focusing on long-context efficiency. Key techniques include KV sharing across layers (Gemma 4), per-layer embeddings (Gemma 4 E2B/E4B), layer-wise attention budgeting (Laguna XS.2), compressed convolutional attention (ZAYA1), and mHC with compressed attention (DeepSeek V4). These reduce KV cache size and memory traffic for reasoning and agent workflows.

Google/DeepMindGoogle/DeepMind PoolsidePoolside DeepSeekDeepSeek Technology Innovation InstituteTechnology Innovation Institute
Sebastian Raschka27.07 · 13:06
Fresh news