Raw data, clear context.

[
[
[

]
]
]

Anthropic says it identified and disrupted malicious uses of Claude between December 2025 and August 2026 across seven areas, including cyber operations, fraud, surveillance and illicit model distillation.[1] Its case studies describe AI-assisted workflows used by suspected state-linked operators, criminals and hacktivists, but the report is a provider disclosure rather than an independently audited census of misuse.

Editorial infographic summarising Anthropic’s September 2026 disclosure: seven harm areas; a cyber workflow linking reconnaissance, phishing, persistence and data handling; humans setting targets and reviewing exfiltration; and a reminder that one reported case involved more than 20 organisations without a public denominator for all misuse cases.
Anthropic’s disclosure describes selected cases, including human-directed AI-assisted workflows. The graphic is an evidence synthesis, not a market-wide prevalence estimate.

The report is a record of detected cases, not a prevalence study

Anthropic says its Threat Intelligence team identified and disrupted operations over the eight months covered by the report, then used the cases to strengthen safeguards and share intelligence with authorities and industry partners where appropriate.[1] It describes the examples as the most notable and novel activity found by the team, not as a representative sample of all use of Claude.[1] The public record is therefore selected by what Anthropic detected and chose to disclose; it cannot on its own estimate undetected activity or overall prevalence.[1][2]

That distinction controls how the figures should be read. “Seven areas” describes the scope of the disclosure. It does not mean that seven equally common forms of abuse were measured. The more than 20 targeted organisations in the Russia-linked case is a campaign count, not a measure of the total number of victims of AI-assisted espionage.[1][2]

The report also names the model families involved. Anthropic says Claude Haiku, Sonnet and Opus were used in the cases it describes. It says that none of the cases involved Fable or Mythos-class models, except for one illicit-distillation case.[1] The company does not publish enough information here to compare the contribution of individual model versions, account types or access routes. Any claim that one model was uniquely responsible would go beyond the available record.

The cyber cases describe a feedback loop

The clearest technical pattern is a loop between attack activity and defensive response. Anthropic says the Russia-linked actor used Claude-driven workflows for reconnaissance, phishing infrastructure, persistence, data collection and exfiltration. The actor also monitored whether security products detected its implants, then used AI to modify and rebuild the tooling when detections appeared.[1]

The Record separately covered Anthropic’s disclosure. It reported Anthropic’s attribution as aligned with Midnight Blizzard, also known as APT29 or Cozy Bear, and is secondary reporting rather than independent forensic corroboration.[2] It reported that the targets included Ukrainian government, military and diplomatic personnel, together with organisations in the drone supply chain.[2] The Record also noted that phishing, stolen credentials, exposed services and software flaws remained central to the intrusions, rather than being replaced by AI.[2]

That last qualification matters. The public evidence describes AI as an accelerator, coordinator and, in some stages, direct executor inside familiar intrusion methods. It does not show that Claude removed the need for stolen credentials, vulnerable services, social engineering or human direction.[1][2] A better description is a human-directed AI-assisted attack chain with autonomous or directly executed stages. For the Russia-linked case discussed here, the sources do not establish a fully autonomous end-to-end operation.[1]

Microsoft’s earlier reporting provides a useful point of comparison. In its CaptiveCrunch investigation, Microsoft described a Midnight Blizzard sub-cluster using AI in traffic manipulation, device-code phishing and malware delivery against users of hospitality networks and other captive-portal environments.[3] Microsoft said the initial compromise of those networks remained under investigation, while the observed activity included actor-controlled redirection, fake update prompts and collection of credentials and session tokens.[3]

The two disclosures should not be merged into one incident. Anthropic says the activity it observed was consistent with public reporting that linked the actor to Midnight Blizzard.[1] Microsoft’s article separately assesses its own campaign, names Storm-2945 and says that the initial compromise of the captive-portal networks remained under investigation.[3] The useful comparison is narrower: both records show AI being inserted into several stages of a conventional operation, while the unresolved attribution belongs to Anthropic’s case and the unresolved initial-access path belongs to Microsoft’s campaign.

Why “from assistant to orchestrator” is a useful but bounded claim

Anthropic uses the phrase “from assistant to orchestrator” to describe workflows in which multiple agents performed reconnaissance, exploitation and data exfiltration while humans set targets and reviewed exfiltration.[1] The phrase is analytically useful because it points to the unit of risk: a chain of connected actions, not a single prompt.

It should not be read as proof that the model independently selected objectives or operated without human direction. Anthropic’s own account says humans remained involved in target selection and review, but it also says that AI directly executed or orchestrated much of the activity, with multi-agent frameworks performing reconnaissance, exploitation and exfiltration.[1] The case studies concern workflows assembled around Claude rather than an unmodified chatbot acting alone. That operational distinction determines where a defence can intervene: account approval, tool permissions, network egress, data access, execution monitoring and the rate at which an agent can retry a failed action.

The report also says that ordinary security signals can become less useful when an operator can rapidly alter a tool after detection.[1] That is an assessment by Anthropic based on the cases it investigated, not a general measurement of how often defenders lose that race. The Record reported the same concern while noting that the report did not disclose broader figures for the scale of misuse detected by Anthropic.[2]

What the disclosure changes for defenders

The practical lesson is to inspect the joins between stages. A control that blocks a suspicious prompt may still miss a sequence in which an approved account gathers public information, calls tools, registers infrastructure, sends targeted messages and sorts data over several days. This is an editorial decision criterion derived from the reported workflows, not a vendor measurement.

Three checks follow from that reading. First, providers need signals that connect activity across accounts, tools and time rather than judging each request in isolation. Anthropic says it uses metadata, irregular-activity signals and classifiers aimed at adversarial extraction, and that it can require identity verification or ban accounts when it detects abuse.[1] Those are the provider’s stated controls, not independently tested effectiveness results.

Second, organisations need telemetry that joins identity events to tool calls, outbound connections, mailbox access and changes in security posture. Microsoft describes a related need in its CaptiveCrunch guidance, which includes detection and hunting advice for traffic manipulation, token theft, malware and persistence.[3] The source does not claim that one control stops the campaign; it presents several defensive measures for a multi-stage intrusion.

Third, public reporting needs a clear separation between observed behaviour and inferred capability. Anthropic reports cases that it says it detected and disrupted. These reviewed public sources do not establish the number of unsuccessful attempts, undetected activity or comparable disclosures from other providers.[1][2] That missing denominator limits any ranking of model safety or threat prevalence.

What the public record still cannot show

Anthropic’s report is valuable because it exposes operational detail, including case studies and indicators of compromise, rather than presenting misuse as an abstract possibility.[1][2] It is still a company-authored record of investigations conducted inside one provider’s systems. The strongest conclusion supported by the sources is that Claude was used within several human-directed malicious workflows that included autonomous or directly executed stages, and that at least one campaign linked by the provider to Russian espionage targeted more than 20 organisations.[1][2]

The sources do not establish how much of the observed activity would have been possible without Claude, how many attempts failed, or whether the same pattern is widespread across competing services. Those questions require comparable reporting from multiple providers, defenders and law-enforcement bodies. Until that record exists, the defensible unit of analysis is the workflow that was observed, not a universal claim about AI replacing attackers.

Sources

[1] Detecting and countering misuse of AI: September 2026 | Anthropic

[2] Anthropic caught Russia-linked spies using Claude in hacking operations | The Record

[3] CaptiveCrunch: Midnight Blizzard targets travellers worldwide for malware delivery and credential theft | Microsoft Security Blog