Where Data Tells the Story
© Voronoi 2026. All rights reserved.

In June 2026, Anthropic told the US Senate that operators tied to Alibaba ran 28.8 million exchanges with Claude through roughly 25,000 fake accounts, an alleged campaign to copy a rival's model by distillation. Distillation is the practice of training a cheaper model on a stronger model's outputs, and it is how much of the open-model ecosystem is built.
The technique is legal when the owner of the teacher model allows it. NVIDIA's Llama-Nemotron post-training set, the largest openly published distillation corpus, was assembled from about 36 million model-generated responses. Of those, 31.5 million (88%) came from Alibaba's own Qwen family, which Alibaba publishes under a permissive license. That single Qwen block is larger than the 28.8 million exchanges Alibaba allegedly harvested from Claude.
So the dividing line is not distillation. It is consent. Distilling Qwen is a free, legal download because Qwen is published; Claude is not, which is reportedly what the alleged campaign's fake accounts were for. The same month, the dispute widened: Apple sued OpenAI over trade secrets, and China moved to curb foreign access to its own open models.
A note on the numbers: the two bars come from different sources and count slightly differently (generated responses versus alleged exchanges), so read them as comparable in scale, not as identical units. The alleged figures are Anthropic's claims, reported by CNBC and Tom's Hardware, and remain unverified.
Sources: NVIDIA Llama-Nemotron dataset card (Hugging Face, July 2026); Anthropic's June 2026 letter to the US Senate.