NewsTradingSentimentCalendarCommunityBriefing
Tech

Microsoft Execs Admit AI Scraping Harms Publishers

By Tech Desk · 2026-09-18 · 3 min read
A stack of printed newspapers with a digital glitch effect distorting the text
Illustration: Tradingbird

Unsealed court documents reveal that Microsoft and OpenAI leadership privately described their data practices as theft and an existential threat to news organizations.

Newly unsealed filings in the copyright lawsuit between The New York Times and two tech giants reveal a stark internal reality: company leaders viewed their AI training methods not as innovation, but as a form of economic disruption. According to documents reviewed by GN technics/ai (en-US), a top Microsoft executive described the mass collection of copyrighted content as the largest theft of labor in human history. This admission contradicts the public stance that such practices are legal fair use.

The lawsuit, filed three years ago, alleges that the companies bypassed paywalls and stripped copyright notices to build their models. While judges have often sided with AI firms on the grounds of fair use, these internal communications suggest that the tech companies knew their actions would directly undermine the financial stability of the publishers whose work they used. The documents paint a picture of an industry aware of the harm it was causing to its own content supply chain.

Internal documents reveal private concerns

The unredacted material includes quotes from senior leaders at both Microsoft and OpenAI that challenge their legal defenses. A Microsoft director of Applied Science described the impact of their products on news sites as a doom loop that hurts both the AI models and the broader web. This internal assessment highlights a paradox where the success of the AI product relies on the decline of the sources that feed it.

OpenAI leadership used equally alarming language in internal communications. The head of ChatGPT described publishers as facing an existential threat from products that are largely substitutive for traditional news consumption. Another executive noted that the models are excellent at news, a capability that directly competes with the original sources rather than merely transforming them. These statements suggest that the companies understood the competitive threat they posed to the media industry.

Data shows severe traffic decline

Microsoft’s own data provides concrete evidence of this disruption. An answer engine within their Copilot product caused click-through rates for The New York Times to drop by as much as 93 percent compared to traditional search results. This significant decline indicates that users are getting their answers directly from the AI, bypassing the original website entirely. The internal documents label this shift as a real risk to the employment of journalists and creators.

Microsoft CEO Satya Nadella testified that paywalled content should be licensed if used for training. He stated that if he had known OpenAI had scraped paywalled information, he would have required the company to retrain its models. This testimony undermines the argument that the use was unintentional or legally protected, suggesting that the companies had the means and the intent to ensure proper licensing but chose not to enforce it.

Legal implications for fair use

The fair use doctrine allows the use of copyrighted work without permission in specific cases, such as criticism or news reporting. However, a key requirement is that the use does not substitute for the original work or harm its market. The internal admissions that AI products are substitutive and threaten the economic foundations of publishers directly conflict with this legal standard. These documents provide strong evidence that the tech companies were aware of the market harm they were causing.

As the legal battle continues, these findings could shift the balance of power in similar copyright cases. The sheer scale of the copying, combined with the explicit internal recognition of the damage, makes it difficult to argue that the practice was benign or transformative. For readers, this means that the AI answers they receive may come from a process that the creators themselves once deemed unethical.

Based on reporting by Yahoo, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories