Moonshot AI

Latest AI news, models and releases from Moonshot AI. ['GLM-5.2', 'Hermes', 'K2', 'K2.6', 'K3', 'K3-mini', 'K3-MoE', 'Kimi', 'Kimi 2.6', 'Kimi-2.6', 'Kimi 3', 'Kimi 3T', 'Kimi Hosted Agent', 'Kimi k1.5', 'Kimi K1.5', 'Kimi K2', 'Kimi K2.5', 'Kimi-K2.5-1T-A32B', 'Kimi K2.6', 'Kimi K2.7', 'Kimi K2.7 Code', 'Kimi K2 Thinking', 'Kimi K3', 'Kimi-K3', 'Kimi K3.5', 'Kimi K5', 'Kimi Linear', 'Kivine', 'Mamba', 'Mamba-2', 'Mamba-3', 'MiniCPM-RobotManip', 'MiniCPM-RobotTrack', 'Minimax', 'Ornith‑1.0‑35B', 'Qwen3-0.6B', 'Qwen3-30B-A3B', 'Qwen3-4B-Instruct-2507', 'qwen3.5-9b', 'Strands Agents']

Agents 🇺🇸

Kimi and kvcache-ai Teams Open-Source AgentENV — a Distributed Platform for Training Agent RL at Kimi K3 Scale

The Moonshot AI team and kvcache-ai have released AgentENV (AENV) under the MIT license — a distributed system for running agent environments using Firecracker micro-VMs. It enables fast sandbox creation, pausing, resuming, and forking, accelerating agent RL training for the Kimi K3 model (2.8 trillion parameters, MoE). The project supports compatibility with E2B, simplifying migration of existing agents.

Moonshot AIMoonshot AI
MarkTechPost28.07 · 00:01
Research 🇺🇸

Managing Reasoning Levels in Large Language Models

The article explains how reasoning models like DeepSeek-R1 and GPT-5.6 learn to produce reasoning traces via reinforcement learning with verifiable rewards (RLVR). It details the training process, inference scaling, think tokens, and how reasoning effort modes (low/medium/high) are implemented through supervised fine-tuning and tokenizer switches.

OpenAIOpenAI DeepSeekDeepSeek Moonshot AIMoonshot AI
Sebastian Raschka27.07 · 20:04
Fresh news