OpenAI reported that it identified and disrupted a coordinated campaign aiming to extract protected reasoning from its models. The company said the earliest activity appeared in the first week of July and that the campaign used systematic, unauthorized techniques consistent with adversarial distillation.
What OpenAI observed
According to OpenAI, operators manipulated model interactions so internal reasoning—information the model records when working through a task—could be reproduced in forms visible to requesters. OpenAI emphasized that attackers did not break encryption, compromise a database, or directly access stored user conversations.
The activity began on July 1 at low volume and peaked on July 24–25, when OpenAI observed about 16,000 requests using a relevant extraction pattern from more than 4,000 users. Further analysis identified related prompt-pattern activity across a cluster of more than 15,000 users; OpenAI said it fully disrupted that activity by July 28. The company also noted that independent security researchers had disclosed related cross-model and conversation-compaction issues and that their reports helped confirm attack paths and accelerate mitigations.
Response and mitigation steps
OpenAI said it responded with account enforcement, technical controls, and partner coordination. Actions included banning or restricting fraudulent accounts, strengthening signup and infrastructure controls, expanding monitoring, and protecting hidden reasoning across users, workspaces, organizations, and model families.
The company reported closing a pathway that allowed replay of encrypted reasoning artifacts, adding checks to detect and hold streamed outputs that might expose reasoning, and working with third-party providers to disrupt related accounts. Findings were shared through the Frontier Model Forum and appropriate government information-sharing channels.
OpenAI attributed a core cluster of the campaign to individuals associated with Moonshot AI, the developer of Kimi. The company noted that the figures it published describe attempted, not necessarily successful, extractions.
Next steps
OpenAI said it expects adversarial distillation attempts to grow more sophisticated and that defending against them will require layered and adaptive controls. The company reported ongoing work to improve tool defenses, classifier coverage, model refusals, and to propagate controls across cloud partners. Mitigation and investigation work is continuing.
Original source: OpenAI News