Tensions between the United States and China over artificial intelligence development have escalated sharply, with top US officials threatening to impose sanctions on leading Chinese AI companies over allegations of large-scale, covert intellectual property theft via a technique called knowledge distillation.
Speaking in an interview with Fox Business, US Treasury Secretary Scott Bessent confirmed that Washington has launched an investigation into whether top Chinese open-weight AI models were developed by illegally extracting proprietary knowledge from American AI systems. “If we find that overseas models are stealing intellectual property from our leading companies, we have the authority to impose sanctions over this illicit activity,” Bessent stated. He added that US investigators have identified digital watermarks from American large language models (LLMs) embedded in multiple Chinese AI models, calling the practice “unacceptable” and saying a final determination on action will come in the coming days or weeks.
The accusations center specifically on Moonshot AI, a prominent Chinese AI developer backed by major domestic technology giants Alibaba, Meituan, and Tencent, which launched its latest flagship model Kimi K3 on July 17. Michael Kratsios, director of the White House Office of Science and Technology Policy, outlined the allegations in a post on X Wednesday, claiming Moonshot AI used knowledge distillation to copy Anthropic’s closed-source Fable model to build Kimi K3. Kratsios alleged the Chinese firm built a custom, sophisticated internal platform to carry out large-scale distillation against US models, using rotating access methods to avoid detection. He also claimed Moonshot AI has acquired GB300 AI servers and accessed the high-performance chips via facilities in Thailand to train its models.
Kratsios emphasized that legitimate, limited use of knowledge distillation — a common AI development technique that allows a smaller “student” model to learn from the outputs of a larger “teacher” model to create more efficient systems — is legal and widely accepted. However, he argued that large-scale covert industrial distillation aimed at stealing proprietary US technology and undermining years of American research investment crosses a clear line.
The US framing of this activity as a national security threat dates back to April, when the White House issued National Security Technology Memorandum 4 (NSTM-4), formally designating “adversarial distillation” as a threat to US national security. The memorandum warned that foreign actors can replicate cutting-edge US AI capabilities at a fraction of the development cost by flooding American public AI interfaces with targeted queries and harvesting the responses, and it directed federal agencies to improve intelligence sharing with private AI companies and explore avenues to hold bad actors accountable.
But Moonshot AI has forcefully denied the allegations. Huang Zhenxin, the firm’s head of business, rejected claims that Kimi K3 relies on distilled data from foreign AI models in comments Tuesday. He attributed Kimi K3’s performance gains to three in-house, original innovations: Moon Clip, a data processing framework that cuts computing costs in half while doubling training efficiency; Kimi Linear Tension, a technique that expands the model’s context window by 10 times; and Attention Residuals, a speed optimization that boosts reasoning performance by 25% that has even drawn public praise from entrepreneur Elon Musk.
The broader dispute has laid bare sharp accusations of double standards from Chinese observers, who point to public examples of US AI developers using the same distillation technique with Chinese open-source models without similar condemnation. A commentary published by Chinese media outlet Guancha.cn argued that US interests frame distillation as innovative progress when American firms do it, but label the exact same practice as IP theft when Chinese developers engage in it.
The commentary highlighted the case of Inkling, the debut AI model from Thinking Machines Lab, a startup founded by former OpenAI CTO Mira Murati. Thinking Machines Lab openly confirmed when launching Inkling in July 2026 that the model’s architecture is based heavily on DeepSeek-V3, a popular Chinese open-source large language model, and its post-training development relied on synthetic data generated by Moonshot AI’s earlier Kimi K2.5 model — a process that falls squarely into the definition of distillation that the US is now threatening to sanction for Chinese firms.
“Chinese open-weight models share their technology freely, lower industry costs, and allow the entire global AI ecosystem to build on their work,” the commentary cited Chinese netizens as saying. “Meanwhile, American closed-source labs hide all their work, charge premium prices, lobby for trade restrictions, and then turn around and accuse everyone else of theft. The hypocrisy is staggering.”
At its core, the current conflict stems from a fundamental divide between two competing AI development models: the open-weight approach embraced by most leading Chinese AI firms, which publishes model weights publicly for global developers to use, modify, and build on, versus the closed-source model dominated by US industry leaders like OpenAI and Anthropic, which keep model weights proprietary and only grant paid access to model outputs via application programming interfaces.
The latest US sanctions threat is the culmination of months of growing tension over the practice. Back in February, OpenAI accused Chinese AI firm DeepSeek of using distillation to free-ride on its frontier AI capabilities, claiming it had detected new covert methods to bypass OpenAI’s access safeguards. Weeks later, Anthropic issued its own accusation, identifying large-scale industrial distillation campaigns run by DeepSeek, Moonshot AI, and MiniMax to illicitly extract capabilities from its Claude model series, using covert tactics to get around access restrictions.
In a June 10 letter to US Senators Tim Scott and Elizabeth Warren, Anthropic detailed more specific allegations against Alibaba, claiming the Chinese tech giant ran a systematic distillation campaign against Anthropic’s Claude models between April 22 and June 5. The letter alleged Alibaba created nearly 25,000 fraudulent accounts to generate more than 28.8 million queries to Claude models, targeting key capabilities including agentic reasoning, software engineering, and long-context tasks. Anthropic urged Congress to improve intelligence sharing between US AI firms, close loopholes that allow Chinese firms to access advanced AI chips, and penalize companies behind what it frames as distillation attacks.
Chinese officials have pushed back against the US accusations. Speaking at the World AI Conference in Shanghai on July 18, Chinese Assistant Foreign Minister Lin Bin did not name the US directly but pushed back against the framing of distillation as a hostile act. “Hype around this issue by some countries is misguided and ultimately counterproductive to global AI development,” he said.
Even within China, some industry commentators have acknowledged that heavy reliance on low-cost distillation carries structural trade-offs for Chinese AI developers. One commentary published on Sina.com noted that during the 2026 FIFA World Cup, users found DeepSeek’s V4-Pro model was unable to answer basic questions about the ongoing tournament, as its training data was frozen in May 2025 and the model generated false explanations rather than acknowledging its knowledge gap. The commentator argued this is an inevitable downside of the low-cost distillation strategy, as the derived models often face higher barriers to updating knowledge and running real-time inference on new information at a reasonable cost.
Notably, the current escalation comes ahead of scheduled high-level AI talks between US and Chinese officials scheduled for September, ahead of a planned meeting between Chinese President Xi Jinping and US President Donald Trump in the US on September 24. Independent benchmark testing of the two models at the center of the current dispute found that Anthropic’s Claude Fable 5 outperforms Moonshot’s Kimi K3 in 22 out of 35 shared evaluation metrics, with particularly strong leads in computer vision and general knowledge tasks. Kimi K3, however, outperforms the US model in long-horizon coding and terminal use benchmarks, and is priced at 70% less than Fable 5: $3 per million input tokens compared to Anthropic’s $10.
