Explainer: What is AI model distillation and why is it becoming a US-China flashpoint?
US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models

A technique that allows developers to shrink powerful artificial intelligence models into cheaper, more efficient systems has become the latest battleground in the intensifying US-China race for AI dominance.
Known as model distillation, the method uses the outputs of a powerful AI system to train a smaller model that can perform some of the same tasks with fewer computing resources.
Long regarded as a standard tool of AI research, distillation is now at the centre of a growing dispute over whether advanced AI capabilities can be transferred without the consent of the companies that created them.
Read: Chinese military researchers tap US AI models to train defence systems
Washington and leading US AI firms have accused Chinese rivals of using the technique to extract capabilities from proprietary models, opening a new front in an increasingly bitter competition over technological leadership.
Below are key facts about the practice at the centre of the debate.
What is model distillation?
The largest AI models, known as frontier models, require enormous amounts of computing power, data and investment to train.
Model distillation offers a way to create smaller systems by using a large "teacher" model to train a smaller "student" model.
The teacher generates examples, such as answers and computer code, which are then used as training material for the student.
The smaller model is not a replica of the teacher. It does not inherit the teacher's weights, architecture or full capabilities. Instead, it learns selected behaviours that enable it to perform specific tasks more efficiently.
Why does distillation matter?
The appeal of distillation is that it can make AI cheaper and easier to deploy.
A frontier model may require large data centres and expensive chips to operate. Distilled models can run on less powerful hardware and be tailored for specific tasks.
That makes them appealing to companies and governments looking to deploy AI more widely, from devices and factories to vehicles and private networks.
Why are reasoning traces important?
Recent AI systems have increased interest in transferring not only final answers but also the steps used to reach them.
These "reasoning traces" can show a smaller model how to approach a difficult problem rather than simply what answer to produce.
Florian Tramèr, an assistant professor at ETH Zurich who researches machine-learning security, compared the process to human learning.
Also Read: EU aims for seven AI gigafactories with €10 billion plan in race with US, China
"If I give you a book of complicated math problems with final solutions, you will have a much harder time learning how to solve problems than if I gave you detailed solutions that describe all steps to take," he said.
As reasoning traces have become more valuable, access to AI outputs has become more sensitive because they may expose some of the methods advanced systems use to tackle complex problems.
Who uses distillation?
Distillation is a widely used AI training technique, not an inherently improper practice.
US researchers and companies have long used it, including Stanford University's Alpaca project and Microsoft's Orca research, which relied on outputs from more advanced models to improve smaller ones.
Chinese researchers have also used outputs from US models in public research projects, including efforts to create Chinese-language instruction models.
The key difference lies in access. Open-weight models let researchers inspect and modify underlying parameters. Closed models, such as OpenAI's ChatGPT and Anthropic's Claude, remain under company control and are typically accessed through proprietary interfaces or APIs.
Why has distillation become a US-China issue?
The controversy is less over distillation itself and more about unauthorised extraction.
AI companies argue there is a distinction between legitimate research and systematically harvesting outputs from proprietary models to replicate commercially valuable capabilities.
Anthropic has accused Chinese entities including DeepSeek, Moonshot and MiniMax of conducting large-scale campaigns to obtain capabilities from Claude models.
Read more: The real threat of AI is not AI
The company said those efforts targeted capabilities including software engineering and advanced reasoning.
OpenAI has also said it has detected attempts by Chinese actors to use its models for distillation-related purposes.
No Chinese companies have accused US rivals of distilling closed-source models so far.


















COMMENTS
Comments are moderated and generally will be posted if they are on-topic and not abusive.
For more information, please see our Comments FAQ