AI distillation moved from a technical debate to a diplomatic flashpoint as Presidents Donald Trump and Xi Jinping prepared to meet at the White House on September 24. U.S. agencies accused Chinese developers of using the technique to extract capabilities from American models; Beijing rejected the charge, while officials explored a narrower area of cooperation on AI-related security incidents.
17
1
39
19
What is AI model distillation?
Distillation trains one AI model using the outputs of another, often more capable model. Developers can use it to build a smaller, less costly system that performs some of the same tasks. It is an established research technique, not a synonym for stealing.
48
The distinction in this dispute is how the training outputs were obtained and used. Distillation within a developer’s own systems, or with permission, differs from using deceptive access to collect a competitor’s outputs in violation of its restrictions. Beijing has emphasized that the technique is widely used, including by American companies; that point does not, on its own, settle the allegations about particular companies’ conduct.
23
5 The available sources do not independently establish how Anthropic uses distillation in its own model development.
What did U.S. agencies allege?
In a September 8 advisory, the FBI, National Security Agency and Cybersecurity and Infrastructure Security Agency warned about what they described as malicious, industrial-scale distillation targeting U.S. AI companies.
1 The agencies named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI and alleged that they extracted billions of tokens across millions of requests to frontier models, including variants of Anthropic’s Claude, OpenAI’s GPT, Google’s Gemini and xAI’s Grok, since at least late 2024.
16
14
Reporting on the advisory describes alleged use of fraudulent accounts and proxy networks to obtain model outputs.
5 The agencies assessed that the activity occurred likely with Chinese government awareness—not that they had established government direction.
16 These remain U.S. allegations, not a finding that every named company committed theft. Separately, Anthropic has accused DeepSeek and Moonshot of routing some user requests to Claude.
49
The issue predated the September advisory: in July, White House science and technology adviser Michael Kratsios accused Moonshot of a large-scale effort to obtain capabilities from leading U.S. models through distillation.
47 The provided reporting does not establish that DeepSeek’s performance or efficiency alone prompted OpenAI’s accusations, and neither would prove unauthorized extraction.
How did China respond?
Chinese officials rejected the allegations of malicious extraction.
49 People’s Daily argued that AI should not be a “monopoly of great powers” and called for cooperation on risk management in a nondiscriminatory development environment.
33 Its response framed the dispute as one about whether U.S. efforts to protect technology also constrain China’s ability to develop AI.
39
That argument and the U.S. allegations address different questions: whether distillation is legitimate in general, and whether specific methods of obtaining proprietary outputs were authorized. The available reporting does not reliably establish a response from each of the six named companies, so a lack of detail here should not be read as silence from any one of them.
Could the rivals cooperate on AI safety?
AI competition was expected to loom over the summit even as officials discussed how to manage shared risks.
17
34 Trump publicly advocated calling AI “super intelligence,” saying that “artificial intelligence” sounded fake, but a change in terminology was not itself a summit policy.
50 The available evidence does not establish the proposed “AI Force” or “AI Czar,” a formal rejection of guardrails, or specific commitments by Xi on human control.
After talks with Chinese Vice Premier He Lifeng in New York, Treasury Secretary Scott Bessent said the United States had proposed a notification mechanism for AI incidents serious enough to affect national security. Trump and Xi were to consider it; the reporting establishes a proposal, not an operational agreement.
18
19
An alert channel could make a dangerous incident easier to communicate about. It would not, by itself, resolve the distillation allegations, set development limits or change U.S. restrictions on advanced AI chips—another point of friction in the relationship.
19
40 That is the summit’s underlying test: whether Washington and Beijing can draw a workable line between legitimate model training and unauthorized extraction while maintaining communication about risks neither can manage alone.
17
48
18