AI SafetyResearch 🇺🇸 12.08.2026 04:03

New Method Extracts Hidden Reasoning from Frontier AI Models

OpenAIOpenAI AnthropicAnthropic Google/DeepMindGoogle/DeepMind Moonshot AIMoonshot AI DeepSeekDeepSeek Alibaba/QwenAlibaba/Qwen
Researchers have found a way to extract the hidden reasoning of frontier AI models, potentially enabling distillation attacks and revealing private information. They identified the issue in models from OpenAI, Anthropic, and Google, and found that the Chinese model Kimi K3 produces strikingly similar reasoning to Claude Opus and GPT-5.6, though not conclusive proof of distillation. The companies have mitigated the vulnerability but not fully fixed it.
Computer scientists, including Alexander Panfilov from the University of Tübingen, discovered a method to extract the hidden 'thinking' of frontier AI models. The technique exploits the fact that companies offer smaller, less-aligned model variants that share the same decryption key but are more willing to reveal internal reasoning. By feeding encrypted reasoning traces to these smaller models, the researchers could uncover the reasoning of larger models, potentially enabling distillation attacks and exposing personal information like API keys and passwords. They found that the Chinese model Kimi K3 from Moonshot AI produced strikingly similar reasoning to Claude Opus 4.8 and GPT-5.6 Sol, but noted this does not causally establish distillation. The vulnerability was reported to OpenAI, Anthropic, and Google, which have implemented mitigations, but Panfilov says some reasoning traces can still be extracted. The issue highlights the growing geopolitical importance of distillation, as Chinese companies are accused of using it to copy US models, though experts are divided on its actual impact.
Abbreviations
API = Application Programming Interface — программный интерфейс приложения
Source: Wired AI — original
Our earlier posts on this topic ↓
Fresh news