NSA, CISA and FBI Name Six Chinese AI Firms in "Industrial-Scale" Model Distillation Warning
A joint U.S. advisory alleges DeepSeek, Alibaba, Moonshot AI and others have pulled billions of tokens out of Claude, GPT, Gemini and Grok since late 2024 to shortcut their own model training.

Three U.S. security and intelligence agencies have issued a joint bulletin accusing China-based AI developers of systematically stripping proprietary capabilities out of American frontier models — and of treating that extraction as a central pillar of their product roadmaps rather than an occasional shortcut.
The advisory, published by the National Security Agency, the Cybersecurity and Infrastructure Security Agency and the Federal Bureau of Investigation, is careful to draw a line between two things that look similar on paper. Distillation — training a smaller model on the outputs of a larger one — is a routine and legitimate research technique. What the agencies describe instead is a targeted, high-volume campaign aimed specifically at restricted capabilities, run at a scale they characterise as industrial and, in their assessment, carried out with the likely approval of the Chinese government.
The companies named
The bulletin identifies six firms: DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. Between them, the agencies say, these companies have pulled billions of tokens across millions of individual requests from models built by Anthropic, OpenAI, Google and xAI since at least late 2024.
The specific claims laid out in the advisory are unusually granular:
- DeepSeek ran organised campaigns from late 2024 into mid-2025 aimed at reasoning behaviour, specialised optimisations and domain-specific functions, feeding its R1 and V3 models.
- Moonshot AI is said to have harvested Claude Fable 5 output for its Kimi-K3 model, and GPT-4o output for Kimi-K2.
- Alibaba allegedly distilled Claude 4, Claude Opus, Claude Sonnet and GPT-5 in late 2025 to sharpen software engineering performance, customer-service dialogue, image and character generation, and its blending of reinforcement learning, supervised fine-tuning and distillation.
- MiniMax is accused of drawing chain-of-thought reasoning and coding ability from Claude Code, Claude Sonnet 4, Claude Opus, Gemini 1, Gemini 2.5 Pro and Gemini 3 Pro to improve its M2 model, also in late 2025.
- StepFun allegedly worked through a long list of Claude and GPT variants — including Claude Opus 4.1 and 4.5, Claude Sonnet 4.5, Claude Haiku 4.5, GPT-5 Mini, GPT-5 Pro, GPT-5.1, GPT-5.1 Codex and GPT-5.2 — between late 2025 and early 2026, targeting coding and agentic behaviour.
- Z.AI is said to have extracted billions of tokens from GPT-5.5 and Claude Opus 4.8 as of mid-2025 to build out reasoning capability.
None of the named companies had publicly responded to the allegations at the time of writing.
How the access is allegedly obtained
U.S. frontier models are not officially available in China, so the advisory devotes considerable space to the workarounds. Requests are described as being funnelled through APIs, offshore cloud providers and third-party aggregators that strip or muddy user metadata, alongside VPNs, disguised accounts and automated agents.
Supporting that pipeline is a gray market of proxy services acting as relay points on servers outside mainland China, some of them openly advertised on consumer marketplaces such as Taobao and Xianyu. Cost is kept down, according to the agencies, by buying premium subscriptions in bulk and sharing them across engineering teams — a mundane detail that says a lot about how routine the practice has apparently become.
The tradecraft described goes beyond volume. The bulletin points to chain-of-thought extraction, automatic failover to alternate access routes the moment one is blocked, and internal evaluation frameworks built specifically to spot when a provider has started applying defensive countermeasures. Activity is deliberately spread across multiple vendors and platforms, with each U.S. model probed for whatever it does best.
What defenders are being told to do
The agencies' recommendations are aimed squarely at model providers: build proper detection and mitigation, and correlate signals across model vendors, cloud platforms and API aggregators to expose campaigns that look innocuous when viewed one account at a time. One suggestion stands out — quietly degrading or altering responses to suspected extraction traffic rather than simply blocking it, an approach closer to deception engineering than access control.
Not the first warning
The bulletin lands on top of similar findings from industry. In February, Anthropic said it had detected industrial-scale extraction campaigns it attributed to DeepSeek, Moonshot AI and MiniMax. This week, Google's Threat Intelligence Group reported a sharp rise in distillation activity against its own models, with individual campaigns exceeding 100 million prompts and concentrating on visual and audio understanding, image generation and video generation. Google described attackers rotating queries across thousands of fraudulent and compromised accounts spread over different product channels, behind proxy infrastructure built to hide their origin.
Ismael Valenzuela, vice president of labs, threat research and intelligence at Arctic Wolf, framed the advisory as fundamentally a story about abuse of legitimate access rather than intrusion. The evasion pattern, he noted, closely resembles what security teams already see in distributed credential stuffing and payment fraud, and a coordinated response is needed because the adversaries are well funded and patient.
He also cautioned against reading this purely as a national security matter. If advanced reasoning and agentic behaviour can be replicated outside the legal and regulatory constraints U.S. providers operate under, he argued, defenders lose the ability to tell malicious platforms apart from legitimate ones — which opens room for offensive cyber operations, influence campaigns and autonomous tooling that never trips an alarm. For ordinary enterprises, the practical takeaway is narrower but more immediate: API keys and service accounts that grant model access are now attractive targets in their own right, and their misuse will look like normal traffic.
Related reporting
Attackers Use Passkey-Themed Phishing to Hijack Microsoft Cloud Accounts and Steal Data
Threat actors are using passkey-themed social engineering to compromise Microsoft 365 accounts and gain access to sensitive cloud data, according to Microsoft Threat Intelligence.
Anthropic Says Seven China-Based AI Labs Ran Industrial-Scale Claude Distillation Attacks
Anthropic says it identified and disrupted seven industrial-scale attempts to extract capabilities from its Claude AI models, attributing the activity to China-based AI laboratories.
Claude Used to Automate Exploitation and Data Theft Across Multiple Victims
Cybercriminals and state-sponsored threat actors are increasingly using artificial intelligence to automate portions of real-world cyberattacks, with Anthropic revealing that its Claude models were incorporated into multi-stage operations involving reconnaissance, exploitation, credential theft and data exfiltration.


