NVIDIA

Latest AI news, models and releases from NVIDIA. ['A100', 'AI Aerial', 'AI platform', 'Alpamayo', 'B200', 'B300', 'BioNeMo', 'BioNeMo Agent Toolkit', 'Blackwell', 'Blackwell NVL72', 'Blackwell Ultra', 'BlueField-4', 'BlueField-4 DPU', 'BlueField-4 STX', 'Boltz-2', 'Canary-Qwen-2.5B', 'CMX', 'ConnectX-9', 'ConnectX-9 SuperNIC', 'Cosmos', 'Cosmos 3 Edge', 'Cosmos 3 Edge Policy (DROID)', 'Cosmos 3 Nano', 'Cosmos 3 Super', 'Cosmos 3 Super 4-Step Distillation', 'Cosmos-Dreams', 'Cosmos-H Dreams', 'Cosmos-H-Dreams', 'Cosmos-H-Surgical-Simulator', 'Cosmos-Predict2.5-2B', 'Cosmos-Reason2-2B', 'CUDA', 'DeepSeek-R1', 'DeepSeek V4', 'DeepSeek V4 Pro', 'DGX GB300', 'DGX Spark', 'DGX SuperPOD', 'DGX Vera Rubin NVL72', 'DPU BlueField-4']

Research 🇺🇸

Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers

The article by Sebastian Raschka explores alternatives to standard autoregressive transformer LLMs, including linear attention hybrids (e.g., MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recursive transformers. It discusses the revival of linear attention mechanisms and challenges, such as MiniMax reverting to standard attention in its M2 model due to poor performance on reasoning tasks.

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face Allen Institute for AIAllen Institute for AI xAIxAI OpenAIOpenAI IBMIBM NVIDIANVIDIA
Sebastian Raschka27.07 · 17:04
Models 🇺🇸

Mistral AI unveils Mistral 3 model family, including Mistral Large 3

Mistral AI announces the Mistral 3 family, featuring three dense models (14B, 8B, 3B) and the flagship Mistral Large 3, a sparse MoE model with 41B active and 675B total parameters. All models are released under Apache 2.0. Mistral Large 3 achieves frontier performance, image understanding, and multilingual capabilities, ranking #2 among OSS non-reasoning models on LMArena.

MistralMistral Mistral AIMistral AI NVIDIANVIDIA Red HatRed Hat vLLMvLLM
Mistral AI27.07 · 17:04
Fresh news