OpenAI has disrupted a coordinated campaign to extract protected reasoning from its models, attributing a core cluster of the activity to individuals linked with Moonshot AI. The campaign involved 16,000 requests from over 4,000 accounts in a 48-hour spike.
TL;DR
- OpenAI disrupted a campaign to extract protected reasoning from its models, attributing it to individuals linked with Moonshot AI.
- The campaign involved 16,000 requests from over 4,000 accounts in a 48-hour spike.
- This incident raises questions about the security of hosted reasoning models and the protection of their internal thought processes.
What happened
OpenAI disclosed on September 30, 2026, that it had identified and disrupted a coordinated campaign designed to extract protected reasoning from its models. The earliest observed activity occurred in the first week of July 2026.
The campaign involved 16,000 requests from more than 4,000 accounts in a 48-hour spike. OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, the Beijing-based company behind the Kimi model family.
OpenAI defined the behavior as adversarial distillation, which involves the systematic and unauthorized use of one model's outputs or reasoning to help train, reproduce, or improve another model.
Why it matters
This incident raises significant concerns about the security of hosted reasoning models and the protection of their internal thought processes. It highlights the challenges faced by AI companies in preventing unauthorized extraction of valuable model capabilities.
For developers building on top of frontier models, this incident underscores the need for robust security measures to protect the reasoning processes of AI models. It also raises questions about the competitive landscape and the potential for similar incidents in the future.
The attribution of the campaign to individuals linked with Moonshot AI adds a layer of complexity to the ongoing debate about AI security and the protection of proprietary model capabilities.
Key facts
- OpenAI disrupted a campaign to extract protected reasoning from its models on September 30, 2026.
- The campaign involved 16,000 requests from over 4,000 accounts in a 48-hour spike.
- OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI.
- The earliest observed activity occurred in the first week of July 2026.
- OpenAI defined the behavior as adversarial distillation.
- The campaign was contained within four weeks of the first spike.
- OpenAI worked with outside researchers and industry partners to build fixes and brief them for feedback.
- OpenAI did not claim that Moonshot AI authorized, funded, or directed the campaign.
Context
This incident is not the first of its kind in the AI industry. Less than two years ago, OpenAI and Microsoft raised similar concerns about DeepSeek, highlighting the recurring pattern of unauthorized extraction of model capabilities.
The incident also underscores the importance of protected reasoning in the competitive landscape of AI. Hidden chain-of-thought reasoning has become a key differentiator for AI models, producing better results on hard reasoning and coding tasks.
The incident raises questions about the security of hosted reasoning models and the potential for similar incidents in the future. It highlights the need for robust security measures to protect the reasoning processes of AI models.
