Meta

Latest AI news, models and releases from Meta. ['Astryx', 'Business Agent', 'BYT5', 'CodeLlama', 'expressive voice AI', 'FAIRChem v2 UMA', 'Gemma 4', 'Gemma4:e2b', 'HuBERT', 'llama', 'Llama', 'LLaMA', 'Llama-2 70B', 'LLaMA2-chat-70B', 'Llama 3', 'Llama 3.1', 'Llama 3.1-405B', 'Llama-3.1-405B', 'LLaMA-3.1 405B', 'Llama 3.1-70B', 'Llama 3.1 8B', 'Llama 3.1-8B', 'Llama-3.1-8B', 'Llama 3.2', 'LLaMA 3.2', 'Llama-3.2-1B-Instruct', 'Llama 3.3', 'llama-3.3-70b', 'Llama 3.3-70B', 'Llama 3 8B', 'Llama3-8B-Instruct', 'Llama 4', 'Llama 4 Maverick', 'LLaMA 65B', 'Llama 70B', 'M2M100', 'M**a AI', 'mBART', 'Meta AI', 'Meta AI model (model name not specified)']

Research 🇺🇸

New DeepSeek-V3 technical report: how hardware-software co-design enables low-cost training of large models

A new technical paper from the DeepSeek team, with CEO Wenfeng Liang as co-author, explores hardware-aware model co-design to reduce LLM training costs. Using DeepSeek-V3 trained on 2048 NVIDIA H800 GPUs as a case study, the paper details innovations in memory efficiency (MLA), sparse computation (DeepSeekMoE), FP8 training, and interconnect-aware routing.

DeepSeekDeepSeek NVIDIANVIDIA Alibaba/QwenAlibaba/Qwen MetaMeta
Synced27.07 · 18:04
Models 🇺🇸

From GPT-2 to gpt-oss: Analyzing Architectural Advances and Comparison with Qwen3

OpenAI released gpt-oss-120b and gpt-oss-20b, their first open-weight models since GPT-2. The architecture features Mixture-of-Experts, Grouped Query Attention, RoPE, SwiGLU, and MXFP4 optimization for local inference. Comparisons with GPT-2 and Qwen3 highlight advances in width vs depth trade-offs and attention sinks.

OpenAIOpenAI Alibaba/QwenAlibaba/Qwen MetaMeta Google/DeepMindGoogle/DeepMind AI21 LabsAI21 Labs TencentTencent
Sebastian Raschka27.07 · 18:03
Fresh news