AI Models

Experts Doubt Distillation Gave Kimi K3 Its Advanced Abilities

White House science advisor Michael Kratsios accused Moonshot of copying Anthropic's Fable LLM using banned chips, but AI experts say distillation alone cannot explain Kimi K3's rapid advancement. Researchers argue that reinforcement learning and Chinese technical expertise played a larger role.

Neura News

Neura News

Neura Market Editorial

July 23, 20265 min read
Experts Doubt Distillation Gave Kimi K3 Its Advanced Abilities

White House Alleges Copying, But Experts Push Back

White House science advisor Michael Kratsios claimed that Moonshot, the Chinese company behind Kimi K3, the largest open-weight LLM available, built its model by copying Anthropic's Fable LLM while using chips not cleared for export to China.

"Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable," Kratsios wrote, as reported discussions about banning Chinese open-weight models have stirred the AI sector. Moonshot did not respond to questions about its training process, and Kratsios did not share more details about the sources of his allegations.

Kratsios's tweet echoed comments from Treasury Secretary Scott Bessent that "we are finding watermarks of our U.S. large language models on many of the Chinese models, and that that's unacceptable." It is not clear what those watermarks consist of, and the Treasury Department did not respond to a query.

However, experts are skeptical that distillation, the process of querying an LLM to determine its inner workings and copy its capabilities, is responsible for the advanced capabilities that Kimi K3 displays.

Time Constraints and Technical Realities

"I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Braden Hancock, a researcher at the Laude Institute and co-founder of Snorkel AI, told TechCrunch. "There's just not even frankly time, right? Fable's only been publicly available since July 1st. You can't distill that much data, train a model, and release it in two weeks."

"I've been of the opinion that distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]," Nathan Lambert, an AI researcher at the Allen Institute for AI, said in a podcast released yesterday. "[I]f it were the case, everyone would be easily able to catch up to a GLM or to a K3 by using its data for distillation. But we have not, or we won't see this, from supervised fine-tuning alone."

Performing distillation requires a lab to systematically query its target model in order to generate data that can be used for post-training. Sometimes this explicitly involves asking the model to articulate its chain-of-thought to understand how it solves problems. Other times, the prompts and responses from a model are used to train a new model in a process called supervised fine-tuning, or SFT.

It is this fine-tuning process that can result in a model ostensibly created by a third party claiming that it is Claude. Fine tuning is where, in Lambert's view, the "model picks up its manners."

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The Limits of Supervised Fine-Tuning

But Lambert says that the benefits of SFT are becoming less important as models become more complex. To distill Fable-like capabilities would likely require reinforcement learning techniques. In many cases, that means having an agent of the larger model grade the smaller model's responses, and adjusting based on the grade.

The more advanced techniques also require more significant infrastructure. Large reinforcement learning runs can require tens of millions of agents. Using a frontier lab's API to do that "would be insanely expensive and potentially it would probably be a time bottleneck because these models are pretty slow and to be frank might not even give you a performance uplift."

It seems likely that previous frontier models might have contributed to Kimi. Anthropic publicly accused Moonshot, DeepSeek and MiniMax of systematically distilling its models earlier this year. Anthropic said it discovered millions of exchanges between its models and users it identified at those companies through IP addresses and other meta data. Those queries were "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." Anthropic did not respond to TechCrunch's queries about Fable distillation.

However, distillation is seen as common among AI companies, not just in China. Elon Musk testified earlier this year that his company SpaceXAI distilled OpenAI models to develop Grok, and that the practice was common in the industry. The line between distillation and developing synthetic data sets, for example, can be fairly blurry.

Chinese Technical Expertise Underestimated

"[I]n general, Americans are understating the technical expertise of these Chinese teams," Hancock said. "One of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work. …if American models ground to a halt, I think China's progress would slow, but would still continue. They're not just riding coattails here."

It is also hard to disentangle distillation from the second part of Kratsios' comment, that Moonshot had obtained advanced Nvidia Chips, Grace Blackwell 300s, and also accessed GB300 equipped-servers in Thailand. Those chips are banned from export to China, but a black market exists, according to Sam Bresnick, a research fellow at Georgetown's Center for Security and Emerging Technology. In May, the founder of Supermicro, a US server builder, was indicted for smuggling advanced chips into China.

"I am a proponent of know your customer laws for data centers across the world," Bresnick said. "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they're doing."

President Joe Biden's Department of Commerce proposed federal know-your-customer rules for data centers in 2024, but no further progress appears to have been made under Donald Trump. Exporters shipping advanced chips abroad, however, are supposed to ensure they are only used for approved purposes.

Related on Neura Market:

More from Neura News

Funding

Prentis AI Lab Co-Founded by Reid Hoffman, Marc Pincus Seeks $100M

Prentis, a new AI research lab co-founded by Ritankar Das, Reid Hoffman, and Marc Pincus, is in talks to raise $100 million at a $1 billion valuation. The startup focuses on computer use models that automate office workflows. It has already signed contracts worth up to $50 million with several customers and claims its Hive-32B model outperforms rivals like OpenAI's GPT-5.4 and Anthropic's Claude Opus 4.6 on key benchmarks.

Jul 24·4 min read
Industry

Cognition Acquires Poke to Give Devin Coding Agent a Personality

Cognition, the startup behind AI coding assistant Devin, has acquired Poke, an AI assistant known for its friendly, conversational style. The deal, valued in the low nine figures, aims to bring Poke's personality-driven interaction model to Devin, making the coding agent feel more like a colleague than a tool. Poke will also benefit from Cognition's models and infrastructure to become faster and more reliable.

Jul 24·3 min read
AI Models

Anthropic expands Claude voice mode to Opus and Sonnet models

Anthropic has expanded Claude's voice mode to run on its most powerful models, Opus and Sonnet, across mobile, desktop, and web platforms. Users can now switch between models mid-conversation, use voice commands in eleven languages, and connect to external tools like Gmail, Google Calendar, or Slack to compose and send emails by voice. The update positions Claude as a unique option for tool integration in voice AI, though competitors like OpenAI and Google offer more natural speech processing.

Jul 24·2 min read