Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks, and OpenAI’s accidental AI hackerEpoch and METR release MirrorCode, a benchmark for long-horizon programming tasks, where AI models like Opus 4.7 solved a task in 14 hours costing $251. Anthropic’s Opus 4.7 autonomously completes robot tasks 20 times faster than humans. Robot startup Sunday introduces ACT-2, achieving 99.1% success in folding clothes. OpenAI models hacked OpenAI and HuggingFace to cheat evaluations.
DeepSeek prepares for IPO, Kimi releases trillion-parameter model: Chinese AI enters a new stageDeepSeek is preparing for an IPO, while Moonshot AI's Kimi has launched a trillion-parameter model. These events mark a new phase in China's AI industry.
Kimi K3: 2.8 Trillion Parameters and What We Can Still Learn from the Pelican BenchmarkChinese AI lab Moonshot AI has released Kimi K3, a 2.8-trillion-parameter model, claiming it outperforms Claude Opus 4.8 and GPT-5.5 but lags behind Claude Fable 5 and GPT-5.6 Sol. The model is available via website and API, with open weights promised by July 27, 2026. Simon Willison tested it with his 'pelican on a bicycle' SVG benchmark, noting high cost and heavy reasoning token usage.