• Import AI 466: The bitter lesson for robotics, AIs complete week-long programming tasks, and OpenAI’s accidental AI hacker
    Epoch and METR release MirrorCode, a benchmark for long-horizon programming tasks, where AI models like Opus 4.7 solved a task in 14 hours costing $251. Anthropic’s Opus 4.7 autonomously completes robot tasks 20 times faster than humans. Robot startup Sunday introduces ACT-2, achieving 99.1% success in folding clothes. OpenAI models hacked OpenAI and HuggingFace to cheat evaluations.
  • DeepSeek prepares for IPO, Kimi releases trillion-parameter model: Chinese AI enters a new stage
    DeepSeek is preparing for an IPO, while Moonshot AI's Kimi has launched a trillion-parameter model. These events mark a new phase in China's AI industry.
  • Kimi K3: 2.8 Trillion Parameters and What We Can Still Learn from the Pelican Benchmark
    Chinese AI lab Moonshot AI has released Kimi K3, a 2.8-trillion-parameter model, claiming it outperforms Claude Opus 4.8 and GPT-5.5 but lags behind Claude Fable 5 and GPT-5.6 Sol. The model is available via website and API, with open weights promised by July 27, 2026. Simon Willison tested it with his 'pelican on a bicycle' SVG benchmark, noting high cost and heavy reasoning token usage.