DeepSeek

Latest AI news, models and releases from DeepSeek. ['DeepSeek', 'DeepSeek 3.1', 'DeepSeek-4-pro', 'DeepSeek API', 'deepseek-coder-v2', 'DeepSeek-GRM', 'DeepSeekMath-V2', 'DeepSeekMoE', 'DeepSeek-Prover-V1.5-Base', 'DeepSeek-Prover-V2', 'DeepSeek-Prover-V2-671B', 'DeepSeek-Prover-V2-7B', 'DeepSeek R1', 'DeepSeek-R1', 'DeepSeek R1-0528', 'DeepSeek-R1-0528', 'DeepSeek-R1-671B', 'DeepSeek-R1-Distilled-Llama-70B', 'DeepSeek-R1-Distilled-Qwen-32B', 'DeepSeek-R1-Distill-Qwen-32B', 'DeepSeek-R1-Zero', 'DeepSeek R2', 'DeepSeek-R2', 'DeepSeek Sparse Attention (DSA)', 'DeepSeek-V2', 'DeepSeek V3', 'DeepSeek-V3', 'DeepSeek V3.1', 'DeepSeek V3.1-Terminus', 'DeepSeek v3.2', 'DeepSeek V3.2', 'DeepSeek-V3.2', 'DeepSeek V3.2-Exp', 'DeepSeek-V3-Base', 'DeepSeek V3/R1', 'DeepSeek v4', 'DeepSeek V4', 'DeepSeek-V4', 'deepseek-v4-flash', 'DeepSeek-V4-Flash']

Research 🇺🇸

Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers

The article by Sebastian Raschka explores alternatives to standard autoregressive transformer LLMs, including linear attention hybrids (e.g., MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recursive transformers. It discusses the revival of linear attention mechanisms and challenges, such as MiniMax reverting to standard attention in its M2 model due to poor performance on reasoning tasks.

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face
Sebastian Raschka27.07 · 17:04
Fresh news