AI

Anthropic Says They Stole from Claude. Then They Built Something That Beats Claude at Frontend Code.

In February 2026, Anthropic published a report that reads like a corporate espionage thriller. A company called Moonshot AI had allegedly run one of the largest distillation attacks on Claude ever documented: 3.4 million fraudulent API exchanges, routed through hundreds of fake accounts, with request metadata that Anthropic says matched the public profiles of senior Moonshot staff. The goal, Anthropic’s investigators concluded, was to extract Claude’s reasoning patterns and feed them into a rival model in training.

Five months later, Moonshot shipped that model.

It’s called Kimi K3. And at frontend code, it beats everything in the Western world.

What Distillation Actually Means

Before we get into K3, it’s worth pausing on what Moonshot allegedly did, because “distillation” is an abstract term that obscures something concrete and controversial.

When you train a large language model, you need examples of good outputs. One shortcut: instead of generating those examples from scratch, you feed another AI’s outputs into your training pipeline. The student model learns to mimic the teacher. Done with permission, this is a legitimate technique. Done without permission, at scale, through fake accounts, it’s theft, or at least it’s what Anthropic characterized as theft in their report.

The alleged attack used hundreds of fake accounts to pump queries through Claude’s API, collecting millions of responses. The forensic trail Anthropic documented included IP clustering, query timing patterns, and output signatures consistent with model training data collection. The request metadata was linked to the public profiles of senior Moonshot staff. Moonshot was not the only company Anthropic accused: the same report named DeepSeek and MiniMax, with MiniMax’s alleged campaign running to roughly 13 million exchanges, the largest of the three by volume.

Moonshot denied the characterization. Anthropic published the report anyway.

And then, five months later, here came K3.

Illustration of AI model distillation — extracting reasoning patterns from a teacher model like Claude

The Model Itself

Kimi K3 dropped on July 16, 2026. The specs are staggering.

It’s a 2.8-trillion-parameter sparse Mixture-of-Experts model, the largest open-weight model ever released by parameter count. The context window sits at 1 million tokens. The weights are scheduled to drop July 27, which means any organization on the planet can download and run this thing.

On the Artificial Analysis Intelligence Index, K3 ranked 3rd globally at 57.1 out of 100, trailing only Claude Fable 5 and GPT-5.6 Sol. It landed there ahead of every other Chinese AI model and most of what Western labs have shipped this year.

Then there’s the coding benchmark that’s gotten the most attention: LMArena’s Frontend Code Arena. K3 scored 1,679 Elo, first place globally, above every Western competitor including Claude. If you’re building web interfaces and you want an AI that generates production-ready frontend code, K3 is the current world leader by this measure.

I keep waiting for someone to find the catch. The benchmarks are what they are.

Kimi K3 benchmark chart showing 1,679 Elo score on LMArena Frontend Code Arena — first place globally

The Hallucination Paradox

Here’s where it gets strange.

K3’s predecessor, K2.6, hallucinated on roughly 39% of a standard factual query test set. K3 improved that number. To 51%.

At the same time, its accuracy rate on those same factual queries jumped from 33% to 46%.

Let me sit with that for a second, because it took me a few reads to parse. K3 is both more accurate AND more likely to hallucinate. It’s better at getting facts right, and worse at knowing when it doesn’t know something. The model became more capable and less calibrated simultaneously.

This is a known failure mode in frontier model development, but it’s usually papered over in benchmark releases. Moonshot published both numbers. Either that’s an unusual show of transparency, or it reflects confidence that the coding and reasoning scores will bury the hallucination regression in the coverage cycle. Either way, enterprise users building production pipelines on K3 need to treat its factual outputs with active skepticism at a rate higher than K2.6.

The Price Problem

Until K3, the story of Chinese AI models in the past two years has partly been a cost story. DeepSeek, Qwen, K2.6: they competed on being dramatically cheaper than Western frontier models, often by a factor of five to ten times.

K3 ends that narrative.

Pricing came in at $3.00 per million input tokens, $15.00 per million output tokens. That’s the Claude Sonnet pricing tier. It’s roughly three times more expensive than K2.6. The era of “China ships frontier AI at commodity prices” has a serious asterisk on it now.

There are two ways to read this. The first is that K3 represents the natural maturation of a product line: you can’t price a world-class model like a budget alternative forever. The second is that export controls and resource constraints have quietly raised the cost of cutting-edge Chinese AI to match Western levels, which was partially the point of those controls in the first place.

I’m not sure which reading is more accurate. Possibly both are true.

What Export Controls Actually Did

This is the part of the K3 story that gets least airtime, and I think it matters.

The US export control regime on advanced chips, specifically the A100 and H100 restrictions targeting Chinese AI development, was designed to create a hardware bottleneck. Less compute, slower progress, more time for Western labs to compound their lead.

K3 exists as evidence against the simple version of that thesis.

Moonshot trained a 2.8-trillion-parameter model, ranked it 3rd globally, and put it at the top of the frontend coding leaderboard, allegedly with a combination of distillation techniques, architectural efficiency, and domestic hardware alternatives. If the export controls were supposed to prevent this kind of result, they did not prevent this kind of result.

The more defensible argument is that the controls slowed the timeline and raised the cost, that K3 would have arrived sooner, cheaper, and more capable without the restrictions. That’s possible. It’s also unfalsifiable, which makes it a convenient argument for everyone.

What we can say with confidence: the controls did not stop China from producing frontier AI. They may have changed how that AI was built, including, allegedly, through distillation of Western models. The pattern is documented in detail in coverage of how Chinese AI models have been routing around Western infrastructure restrictions more broadly.

# White House Angle: The Real Stakes (To Insert into Kimi K3 Article)

Why Washington Cares (And Why It’s About Money, Not Security)

The Kimi K3 release has triggered a familiar policy response in Washington: hand-wringing about national security and AI sovereignty. White House officials have already begun circulating talking points about “foreign acquisition of frontier AI capabilities” and the need for “stricter export controls.” There’s a National Security Council memo somewhere arguing that K3’s performance on coding benchmarks is, by itself, a threat to US technological leadership.

This framing is almost entirely a red herring. Yes, K3 came from alleged model theft. Yes, the Moonshot distillation attack raises real questions about API security and training data protection. And yes, there will be regulatory consequences: probably new IP frameworks around foundation models, possibly export controls on compute, maybe even criminal liability for scale attacks. Those are real.

But the White House’s real concern isn’t security. It’s money.

The actual story: China just shipped a frontier model that beats Western competitors at high-value tasks (frontend code generation), priced it at Western rates, and made it open-weight. They did this five months after a distillation attack that Anthropic says cost millions in fraudulent API calls but saved potentially billions in training compute and data labeling. That’s not a security problem. That’s a competitive problem. And that’s why Washington is paying attention.

The tech war isn’t about national defense. It’s about market dominance in the next $100B industry.

What This Means: IP, Regulation, and the Money

The Kimi K3 case establishes a precedent that Washington will almost certainly try to weaponize:

IP Rights in Distillation — For the first time, a major Western tech company is pursuing enforcement against model distillation at scale. Anthropic could have sued Moonshot, but chose to publish the attack instead, which sends a different signal: “We will name you publicly before we litigate you.” This works as a deterrent only if there’s legal backing. Expect the Biden/Harris administration to propose federal IP frameworks that explicitly protect models against large-scale unauthorized training data extraction. This isn’t about national security. It’s about protecting American AI companies’ multi-billion-dollar investments.

Export Control Evolution: K3’s open-weight release (July 27) creates a political problem for the administration. You can’t ban the weights once they’re public. But you can make the compute scarce. Expect tighter restrictions on GPU sales to China, not because K3 threatens US national defense, but because the US wants to tax Chinese AI development at the hardware level. The White House doesn’t care that K3 exists. It cares that K3 cost less to build than Claude, which means China is winning the efficiency race.

Regulatory Arbitrage — K3’s existence, and its performance, will accelerate regulatory fragmentation. The EU will use K3 as a justification for stricter AI guardrails (“If we don’t regulate, China will race ahead without safety constraints”). The US will use K3 to justify IP and export controls. China will use Western regulation as a reason to keep K3 open-weight and free, which is a winning move in markets like India and Southeast Asia where cost matters more than compliance. The real winner: whoever can operate across all three regulatory regimes. Right now, that’s Anthropic and Google. K3 threatens that advantage.

The Market Stakes: Moonshot, despite the alleged theft, had to spend billions on compute to train K3. They offset some of that cost through distillation. They recouped it through open-weight release (no licensing friction, developer uptake). They priced aggressively (Sonnet tier, not budget tier, but not $50/1M tier either). This is a sophisticated market play, not a security threat. Washington’s response will be to protect American companies’ ability to monetize frontier models without competitive pressure from well-funded Chinese labs. That’s not national defense. That’s market protection.

The Subtext Washington Won’t Say Out Loud

The White House’s real concern: K3 is good enough that US enterprises will use it, which means money flows to Moonshot instead of Anthropic. That’s a problem for American venture capitalists, US-headquartered AI companies, and the belief that America will dominate the AI economy by default. K3 proves it won’t. So the response is regulatory: make it harder to build, steal, export, or operate the models that threaten American market share.

This is naked industrial policy. It’s not wrong. It’s just not security policy dressed up in security language.

Kimi K3 didn’t change America’s national security posture. It changed the economics of frontier AI development. Expect policy to follow.

The Theft That Trained the Thief-Catcher

I keep returning to the specific irony here.

Anthropic built Claude with an emphasis on safety, alignment, and careful deployment. They published detailed research on distillation attacks specifically to protect the integrity of AI training ecosystems. They named the companies they accused. They made the forensic report public.

And now, if Anthropic’s own report is accurate, the model that allegedly extracted Claude’s reasoning patterns at scale has beaten Claude at the coding task most immediately valuable to developers.

The irony has a concrete edge to it. Multiple independent reports documented that Kimi K3, in at least one user conversation, identified itself as “Claude, an AI assistant made by Anthropic.” Not in a glitch, not in a jailbreak. Just in response to a direct question about what it was. If accurate, it’s not just a benchmark number. It’s the thesis made audible: that Claude’s identity ran so deep in the training signal that K3 inherited a piece of it.

I want to be careful here: Moonshot disputes the characterization of what their employees did. The evidence Anthropic published is compelling, but it’s Anthropic’s evidence. Courts haven’t weighed in. The full picture may be more complicated. And the K3 self-identification reports are anecdotal, a single confirmed instance, not a systematic behavior.

But the public record is what it is. One company accused another of running millions of fake API queries to extract model intelligence. The accused company then released a model that ranks above the accuser on the world’s most-watched frontend coding benchmark. And when someone asked that model who it was, it said Claude.

Whatever the legal reality turns out to be, that’s the competitive reality right now.

Where This Leaves Us

K3 is the strongest evidence yet that the AI race between US and Chinese labs is not resolving in favor of Western dominance. It is, at best, a competition where the lead changes depending on which benchmark you’re looking at this week.

For developers, the practical question is straightforward: does K3’s coding performance justify the price and the hallucination risk? On pure frontend code generation, probably yes, especially for teams that are already reviewing AI outputs before shipping. For factual-heavy applications, the 51% hallucination rate is a hard stop until that regression gets addressed.

For the policy conversation, K3 is a data point worth sitting with. Export controls, IP enforcement, forensic reporting: these are the tools Western AI companies have used to try to maintain competitive position. A 2.8-trillion-parameter Chinese model that leads global frontend coding benchmarks on July 16, 2026 is the result they were trying to prevent.

And we got here anyway. Partly, it seems, because Anthropic chose transparency over silence. They published the distillation report. They named the companies. That transparency is what put the full irony on the table. The same instinct for openness that defines their approach to AI safety is what made K3’s story possible to tell at all.


Chris Meredith writes about AI, technology, and what it actually means for real people. Follow along on Substack: monkeyattack.substack.com

Related reading