Moonshot AI

Latest AI news, models and releases from Moonshot AI. ['GLM 5.2', 'GLM-5.2', 'Hermes', 'K2', 'K2.6', 'K3', 'K3-mini', 'K3-MoE', 'Kimi', 'Kimi 2.6', 'Kimi-2.6', 'Kimi 3', 'Kimi 3T', 'Kimi Hosted Agent', 'Kimi k1.5', 'Kimi K1.5', 'kimi-k2', 'Kimi K2', 'kimi-k2-0711-preview', 'Kimi K2.5', 'Kimi-K2.5-1T-A32B', 'Kimi K2.6', 'Kimi K2.7', 'Kimi K2.7 Code', 'Kimi K2 Thinking', 'Kimi K3', 'Kimi-K3', 'Kimi K3.5', 'Kimi K4', 'Kimi K5', 'Kimi Linear', 'Kivine', 'Krea 2', 'Krea 2 Turbo', 'Mamba', 'Mamba-2', 'Mamba-3', 'MiniCPM-RobotManip', 'MiniCPM-RobotTrack', 'Minimax']

Research 🇺🇸

Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers

The article by Sebastian Raschka explores alternatives to standard autoregressive transformer LLMs, including linear attention hybrids (e.g., MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recursive transformers. It discusses the revival of linear attention mechanisms and challenges, such as MiniMax reverting to standard attention in its M2 model due to poor performance on reasoning tasks.

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face
Sebastian Raschka27.07 · 17:04
Research 🇺🇸

State of LLMs in 2025: Progress, Challenges, and Predictions

In 2025, LLM development was dominated by reasoning models using reinforcement learning with verifiable rewards (RLVR) and the GRPO algorithm, sparked by DeepSeek R1. Key trends include a focus on inference-time scaling, MoE architectures, and efficiency tweaks like Gated DeltaNets. Academic research highlighted GRPO variants with improvements such as zero gradient filtering and token-level loss.

DeepSeekDeepSeek OpenAIOpenAI Alibaba/QwenAlibaba/Qwen Moonshot AIMoonshot AI NVIDIANVIDIA Google/DeepMindGoogle/DeepMind
Sebastian Raschka27.07 · 17:03
Fresh news