Meta

Latest AI news, models and releases from Meta. ['Astryx', 'Business Agent', 'BYT5', 'CodeLlama', 'expressive voice AI', 'FAIRChem v2 UMA', 'Gemma 4', 'Gemma4:e2b', 'HuBERT', 'llama', 'Llama', 'LLaMA', 'Llama-2 70B', 'LLaMA2-chat-70B', 'Llama 3', 'Llama 3.1', 'Llama 3.1-405B', 'Llama-3.1-405B', 'LLaMA-3.1 405B', 'Llama 3.1-70B', 'Llama-3.1-70B', 'Llama 3.1 8B', 'Llama 3.1-8B', 'Llama-3.1-8B', 'Llama 3.2', 'LLaMA 3.2', 'Llama-3.2-1B-Instruct', 'Llama 3.3', 'llama-3.3-70b', 'Llama 3.3-70B', 'Llama 3 8B', 'Llama3-8B-Instruct', 'Llama 4', 'Llama 4 Maverick', 'LLaMA 65B', 'Llama 70B', 'M2M100', 'M**a AI', 'mBART', 'Meta AI']

Research 🇺🇸

Beyond Standard LLMs: Linear Attention Hybrids, Text Diffusion, Code World Models, and Small Recursive Transformers

The article by Sebastian Raschka explores alternatives to standard autoregressive transformer LLMs, including linear attention hybrids (e.g., MiniMax-M1, Qwen3-Next, DeepSeek V3.2, Kimi Linear), text diffusion models, code world models, and small recursive transformers. It discusses the revival of linear attention mechanisms and challenges, such as MiniMax reverting to standard attention in its M2 model due to poor performance on reasoning tasks.

Moonshot AIMoonshot AI MiniMaxMiniMax Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek Google/DeepMindGoogle/DeepMind MistralMistral MetaMeta Hugging FaceHugging Face
Sebastian Raschka27.07 · 17:04
Agents 🇺🇸

How to set up a local AI agent for coding and avoid dependency on rising cloud model prices

The Register guides readers on setting up a local AI coding assistant using the Qwen3.6-27B model to avoid rising costs of cloud APIs. The article covers inference setup, hyperparameters, and three agent frameworks: Claude Code, Pi Coding Agent, and Cline. It concludes that while local models are not yet replacements for frontier models, they are surprisingly capable for many tasks.

Alibaba/QwenAlibaba/Qwen AnthropicAnthropic OpenAIOpenAI MicrosoftMicrosoft MetaMeta
The Register AI/ML27.07 · 16:03
Fresh news