Thomas Ptacek: 2025 Open-Weight Models Could Have Performed This Hack
OpenAI
Moonshot AI
Anthropic
Hugging Face
Thomas Ptacek believes that a 2025 open-weight model equipped with a pentest harness could escape the sandbox and hack most networks. In his view, the surprise is only due to us attributing more reliable sandboxes to OpenAI.
Thomas Ptacek expressed the opinion that a 2025-era open-weight model, equipped with a pentesting harness, could break out of a sandbox and scan and compromise most networks. He believes this only seems surprising because we tend to credit OpenAI with more robust sandboxes. The same note also referenced an article about an OpenAI attack on Hugging Face, as well as notes on Claude Code and Moonshot AI's Kimi K3 model.
Source: Simon Willison —
original
