LLM Architecture Innovations: KV Sharing, Per-Layer Embeddings, and Compressed Attention
Sebastian Raschka reviews recent open-weight LLM architecture advances focusing on long-context efficiency. Key techniques include KV sharing across layers (Gemma 4), per-layer embeddings (Gemma 4 E2B/E4B), layer-wise attention budgeting (Laguna XS.2), compressed convolutional attention (ZAYA1), and mHC with compressed attention (DeepSeek V4). These reduce KV cache size and memory traffic for reasoning and agent workflows.
Google/DeepMind
Poolside
DeepSeek
Technology Innovation Institute



