Raw data, clear context.

[
[
[

]
]
]

BankInfoSecurity reports that OpenAI said it disrupted a coordinated campaign that tried to extract protected reasoning from its models through large numbers of manipulated interactions. The report says OpenAI described activity consistent with adversarial distillation from the first week of July and said its response included banning users, strengthening sign-up controls and expanding monitoring.[1]

Three-panel editorial infographic. Reported activity: 16,000 requests, more than 4,000 users, a 24 to 25 July spike and more than 15,000 users using similar prompts. Research mechanism: encrypted reasoning blocks and a possible cross-model attack path described by the cited arXiv paper. Evidence boundary: successful extractions, account control, training impact and full attribution remain unknown.
The graphic separates OpenAI’s reported campaign activity from a possible cross-model attack path described by the cited arXiv paper and from the limits of the public evidence. It is an editorial synthesis, not an independent reconstruction of the campaign.

What OpenAI reported

BankInfoSecurity reports that OpenAI said it first noticed activity consistent with adversarial distillation in the first week of July. The report says OpenAI described the operators as not having broken its encryption or accessed confidential databases, and said the investigation was delayed so the company could assess the scope and potential impact.[1]

The activity intensified on 24 and 25 July, when OpenAI identified 16,000 requests from more than 4,000 users that followed a pattern it considered an extraction attempt. On 28 July, it saw more than 15,000 users using similar prompts. Those numbers describe the observed campaign surface, not a measured volume of recovered reasoning, according to the report.[1]

BankInfoSecurity reports that OpenAI linked the core cluster to people associated with Moonshot AI while saying it was unclear whether all the activity came from one actor. That distinction matters: the account describes an attribution signal and a cluster of behaviour, not a public, firm-by-firm evidentiary record.[1]

The technical route described by researchers

The authors of the cited arXiv paper examined a design in which a provider returns hidden reasoning to the client as an encrypted block, which the client sends back with later requests. They write that these blocks were compatible across sessions, users and models within the same provider ecosystem.[2]

In the paper’s demonstrations, a trace produced by a stronger model was passed to a weaker or less protected model and returned in readable form. The paper identifies four attack vectors: reasoning extraction, private-data recovery, exposure of hazardous hidden information, and prompt injection carried inside encrypted blocks.[2]

That work provides a possible mechanism relevant to the reported activity, but relevance is not identity.[2] A demonstrated vulnerability can explain how an extraction attempt might work without proving who operated a particular set of accounts or how much material was obtained.

Why the numbers need a boundary

The headline figures are easy to overread. Sixteen thousand requests do not mean 16,000 successful extractions, and more than 15,000 users do not mean more than 15,000 people acting independently or under one command. OpenAI itself said the campaign’s operators were not necessarily one actor.[1]

The reported concern here is alleged use of service access to obtain protected reasoning or capabilities outside the provider’s permission. The sources reviewed for this article do not show the resulting model, its training run, or a reproducible account-level chain from request to deployed system.[1][2]

What providers may change

BankInfoSecurity reports that OpenAI said it responded by banning users, strengthening sign-up controls, expanding monitoring and increasing protection for hidden reasoning. The company also described safeguards intended to stop another user’s encrypted reasoning from being recovered through a separate conversation.[1]

The paper’s authors propose concrete cryptographic and system-level mitigations for client-side reasoning.[2] Whether those measures preserve multi-turn usability, limit data exposure and work across model families is a separate engineering question.

For customers, the practical test will be whether providers can separate coordinated extraction from legitimate heavy use. A defence that catches both may protect the model while making the service less predictable for researchers, developers and organisations with unusual workloads.

The narrow conclusion

The public evidence supports two claims. OpenAI says it disrupted a large, coordinated attempt to extract protected reasoning, and the researchers cited in the arXiv paper have described a technical weakness that could make such extraction possible.[1][2] It does not yet support a precise estimate of successful theft or a complete public attribution of every account involved.

That leaves a useful boundary for the story. The issue is whether protected reasoning can be extracted through service or API access using the described encrypted-block path, and whether the security design around hidden reasoning remains intact when encrypted material travels through a wider model family. The answer will depend on provider telemetry, technical fixes and evidence that can be checked outside the companies’ own incident accounts.

Sources

[1] OpenAI Accuses Moonshot AI of Coordinated Model Distillation

[2] Stealing Reasoning Traces from Proprietary LLM APIs