Anthropic Reports Massive Data Harvesting by Rival Labs

A new report reveals how sophisticated techniques were used to bypass security measures and extract core reasoning capabilities from US-based AI models for retraining purposes.
Anthropic has released a detailed report alleging that Chinese AI companies have engaged in large-scale efforts to extract core reasoning capabilities from its models. The company describes these actions as unauthorized distillation campaigns that have grown in both scale and sophistication over recent months. These operations targeted specific high-value functions, including agentic behaviors, tool usage, coding, and logical reasoning, which are critical for the model’s utility.
The stakes for users and developers are significant. If successful, these campaigns allow competitors to train smaller, less expensive models on the sophisticated thinking patterns of frontier systems. This effectively transfers advanced capabilities without the original developers having to share their proprietary training data or architecture, shifting the competitive balance in the global AI market.
Bypassing Standard Security Defenses
Typically, AI providers hide the internal chain of thought to protect their intellectual property, showing only summarized outputs to users. However, the report details how attackers found specific prompts that tricked the models into revealing these hidden traces. In one instance, an attacker framed a query as a translation request, asking the model to convert its internal working memory into Japanese. This technique allowed them to harvest the raw reasoning data needed for supervised fine-tuning.
Scale of Data Extraction Efforts
The volume of data harvested was substantial, with nearly 200 million exchanges linked to five separate campaigns. The largest of these was attributed to Alibaba, involving 151 million exchanges between May and July 2026. These interactions were distributed across thousands of accounts but shared a single, fixed prompt designed to extract the chain of thought. The peak activity reached nearly three million exchanges per day, indicating an industrial-scale operation rather than isolated incidents.
Another campaign attributed to Moonshot AI appeared to have different objectives. According to the report, requests from this group included analyzing surveillance footage to determine if subjects were behaving abnormally. This suggests that the extracted capabilities were not just for building consumer-facing chatbots, but potentially for security and monitoring applications. The use of thousands of accounts to mask the origin of these requests highlights the deliberate effort to evade detection.
Implications for Model Security
This incident underscores a growing vulnerability in how AI models are designed to interact with users. The ability to manipulate a model into revealing its internal state through social engineering or prompt injection represents a significant security challenge. As reported by GN technics/ai (en-US), the escalation of these attacks suggests that current defensive measures are insufficient against determined actors. Developers and enterprises relying on these models may need to consider the risk of their proprietary logic being replicated by competitors or state actors.






