NewsTradingSentimentCalendarCommunityBriefing
Tech

Microsoft Exec Warned of Massive Labor Theft in AI Data Collection

By Tech Desk · 2026-09-17 · 2 min read
A tangled ball of loose wires and cables
Illustration: Tradingbird

Internal documents reveal that senior Microsoft officials privately feared their AI training methods constituted a historic violation of creator rights.

Long-standing efforts by Microsoft and OpenAI to keep sensitive details confidential are unravelling. Newly unsealed internal documents show that executives inside Microsoft were deeply concerned about the legal and ethical implications of using news content to train their AI models. These concerns emerged well before the launch of popular products like ChatGPT and Copilot, suggesting that the company was aware of the potential for significant conflict with publishers.

The leaked materials include sharp warnings from Brent Hecht, a Director of Applied Science at Microsoft. In internal communications, he described the practice of scraping news for AI training as an astonishing theft of unprecedented proportions. Hecht went so far as to call it perhaps the largest theft of labor in human history, a phrase that starkly contrasts with the public narrative of innovation and progress often associated with these technologies.

Internal doubts about fair use claims

While the companies have publicly argued that training AI on existing content falls under fair use, internal records paint a different picture. Hecht’s documents suggest that the scale of data collection made a complete mockery of the fair use concept. This disconnect between private skepticism and public legal strategy could complicate ongoing litigation led by major news organizations, who are seeking to hold the AI firms accountable for copyright violations.

The risk was not just legal but structural. One document described a doom loop where the threat to the economic foundations of content suppliers would eventually hurt the performance of the AI models themselves. This indicates that leadership understood the dependency of their products on a healthy web ecosystem, even as they pursued aggressive data acquisition strategies.

Existential threat to news publishers

Nick Turley, the head of ChatGPT at OpenAI, also contributed to the internal discourse. In a message cited by plaintiffs, he acknowledged that publishers would face an existential threat from commercial products trained on their content. These tools are designed to substitute for traditional news providers, creating a direct economic conflict. The recognition of this threat within the company hierarchy underscores the magnitude of the disruption these technologies pose to the media industry.

The situation highlights a fundamental tension in the development of large language models. The companies rely on vast amounts of human-created content to train their systems, yet the resulting products compete with the creators of that content. As these internal documents become part of the public record, the argument that this is a benign or transformative use of data faces its strongest challenge yet.

Implications for the content supply chain

A Microsoft document explicitly noted the unusual nature of an end-product threatening the economic foundations of its essential suppliers. This self-inflicted crisis in the content supply chain raises questions about the long-term sustainability of the current AI development model. If the sources of training data are weakened or shut down, the quality and legality of future models could be severely impacted.

The exposure of these documents by GN technics/ai (en-US) and other outlets marks a turning point in the debate over AI and copyright. It shifts the conversation from abstract principles to specific internal admissions of risk and wrongdoing. For readers, the stakes are clear: the future of online information depends on whether companies can build AI systems that respect the labor and rights of those who created the digital world.

Based on reporting by Ars Technica, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories