DeepSeek

Latest AI news, models and releases from DeepSeek. ['DeepSeek', 'DeepSeek 3.1', 'DeepSeek-4-pro', 'DeepSeek API', 'deepseek-coder-v2', 'DeepSeek-Coder-V2-Lite', 'DeepSeek-GRM', 'DeepSeekMath-V2', 'DeepSeekMoE', 'DeepSeek-Prover-V1.5-Base', 'DeepSeek-Prover-V2', 'DeepSeek-Prover-V2-671B', 'DeepSeek-Prover-V2-7B', 'DeepSeek R1', 'DeepSeek-R1', 'DeepSeek R1-0528', 'DeepSeek-R1-0528', 'DeepSeek-R1-671B', 'DeepSeek-R1-Distilled-Llama-70B', 'DeepSeek-R1-Distilled-Qwen-32B', 'DeepSeek-R1-Distill-Qwen-32B', 'DeepSeek-R1-Zero', 'DeepSeek R2', 'DeepSeek-R2', 'DeepSeek Sparse Attention (DSA)', 'DeepSeek-V2', 'DeepSeek-V2.5', 'DeepSeek V3', 'DeepSeek-V3', 'DeepSeek V3.1', 'DeepSeek V3.1-Terminus', 'DeepSeek v3.2', 'DeepSeek V3.2', 'DeepSeek-V3.2', 'DeepSeek V3.2-Exp', 'DeepSeek-V3-Base', 'DeepSeek V3/R1', 'DeepSeek v4', 'DeepSeek V4', 'DeepSeek-V4']

Research 🇺🇸

LLM Architecture Innovations: KV Sharing, Per-Layer Embeddings, and Compressed Attention

Sebastian Raschka reviews recent open-weight LLM architecture advances focusing on long-context efficiency. Key techniques include KV sharing across layers (Gemma 4), per-layer embeddings (Gemma 4 E2B/E4B), layer-wise attention budgeting (Laguna XS.2), compressed convolutional attention (ZAYA1), and mHC with compressed attention (DeepSeek V4). These reduce KV cache size and memory traffic for reasoning and agent workflows.

Google/DeepMindGoogle/DeepMind PoolsidePoolside DeepSeekDeepSeek Technology Innovation InstituteTechnology Innovation Institute
Sebastian Raschka27.07 · 13:06
Agents 🇷🇺

How to train an infrastructure LLM agent not to lie

A read-only AI agent for infrastructure incident investigations was built. Initially it gave plausible but useless answers. The team solved the problem by moving LLM inside the application, adding tool contracts, evidence guards, and structured planning. In the last run, 44 of 45 test scenarios passed.

DeepSeekDeepSeek
Habr — хаб ИИ27.07 · 13:02
Fresh news