Chinese military AI distillation uses U.S. model outputs

Facebook
Twitter
LinkedIn
Pinterest
Pocket
WhatsApp
Chinese military AI distillation uses U.S. model outputs

Chinese military AI distillation is drawing on outputs from leading U.S. systems built by OpenAI and Anthropic to train homegrown models that support defense development, according to a review of more than 80 Chinese academic papers and patents.

The previously unreported body of work provides an unusual view into how security and military-affiliated organizations in China are using advanced U.S. artificial intelligence as a shortcut to build specialized domestic tools, despite U.S. efforts to curb Beijing’s access to cutting-edge chips and other strategic technologies.

The documents detail broad adoption of a method known as model distillation, in which responses from a powerful AI system are used to teach smaller, task-focused models that can be operated locally without the vast computing needed to create frontier-scale systems from the ground up.

The review, which included research compiled by the Washington-based Jamestown Foundation and shared with the reviewers, found that distillation is commonly applied by scholars tied to the People’s Liberation Army and other military bodies.

The papers indicate that Chinese defense entities view top U.S. AI systems both as a technical reference and as a means to narrow the performance gap with American counterparts.

The dispute turns on alleged unauthorized extraction rather than the concept of distillation itself, which is widely practiced in the industry.

The issue has become a focal point ahead of planned U.S.-China discussions on AI safety and governance. U.S. officials have alleged that some Chinese groups use distillation to pull capabilities from American models, potentially undercutting export controls and violating intellectual property rights.

China has rejected those claims, accusing Washington of pursuing AI hegemonism and arguing that U.S. firms have engaged in comparable methods.

Chinese developers have also pushed back on assertions that their progress depends on foreign systems. Last week, AI startup Moonshot denied allegations by the Trump administration that its Kimi K3 model relied on distillation, saying it resulted from proprietary advances.

Sunny Cheung, a Jamestown Foundation fellow who examined more than 60 of the papers, said Chinese military researchers are methodically capturing the reasoning processes of Western models to tailor them for surveillance, cyber operations and battlefield decision support.

“Teaching a model the right answer is one thing but teaching it the reasoning behind the answer is much harder,” Cheung said, adding that the studies suggest an effort to transfer costly, proprietary reasoning into smaller, locally controlled systems.

The reviewers verified the academic literature and identified roughly two dozen additional case studies with military links.

One paper published last year by researchers in PLA Unit 96941, an intelligence and cyber-warfare unit in Beijing, described using OpenAI’s GPT-3.5 to process sensitive military source code.

The researchers wrote that third-party models were not suitable for handling classified data. To address that concern, they used GPT-3.5 to summarize software code, then trained a domestic model on those summaries so it could run fully within Chinese military networks.

The White House, the Department of Defense, China’s Ministry of Foreign Affairs, the PLA and OpenAI did not respond to requests for comment.

Chinese military AI distillation in practice

Reuters’ and Jamestown’s review found Chinese researchers applying distillation to tasks ranging from content oversight to defense deployment.

At the North University of China, which maintains close ties to the nation’s weapons sector, researchers used Anthropic’s Claude 3 Haiku to create synthetic training data for a text classification system aimed at social media monitoring and moderation.

Anthropic said it does not provide commercial access to Claude in China or to firms controlled by Beijing and uses monitoring systems to detect policy violations.

The company also warned that distilled models can shed the original systems’ safety guardrails, which could allow sensitive capabilities to migrate into models outside its control.

A 2024 paper from the PLA’s National University of Defense Technology detailed shrinking an image-processing model via distillation for onboard use in unmanned aerial vehicles, enabling drones to analyze live video and assist with navigation and targeting in real time even without communications links.

In a separate study published earlier this year, researchers at the Academy of Military Sciences used distillation to run a target-recognition model on tactical hardware during simulated maritime missions involving drones, surface vessels and unmanned submarines.

From code analysis to battlefield applications

China has leaned into distillation as it competes with the United States in frontier AI while facing constraints on advanced computing resources due to U.S. export controls on high-end chips.

Central and local authorities have promoted model lightweighting and edge computing, channeling subsidies and research funds into technologies that allow AI to run on drones, satellites and other platforms with limited processing power.

Experts caution that distillation carries important trade-offs.

As domestic models improve, Chinese military researchers are also studying distillation as a potential security vulnerability.

In January, scholars at the Army Engineering University published findings on the risks of data-free distillation, a technique for inferring a model’s abilities without direct access to its parameters.

They proposed defenses intended to obscure the hidden logical information that can be exposed through a model’s publicly visible outputs.

Distilled systems inherit specific skills but do not fully reproduce the broad capabilities of frontier models.

Trevor Koverko, co-founder of AI data company Sapien, said distilled models remain less capable than their teacher models.

“It is best understood as transferring selected capabilities into a cheaper, locally controlled system, not achieving independence from frontier AI,” Koverko said.

The use of distilled AI in unmanned systems and targeting echoes broader trends in modern warfare, including programs such as the Pentagon GBAM challenge that seek cheaper long-range strike capabilities.

The Jamestown Foundation, a Washington-based research institute, continues to publish analyses on China’s military and technological development at its official website.

Facebook
Twitter
LinkedIn
Pinterest
Pocket
WhatsApp