NewsTradingSentimentCalendarCommunityBriefing
Tech

Internal Docs Reveal AI Giants Warned of Web Collapse

By Tech Desk · 2026-09-18 · 3 min read
A tangled knot of digital threads unraveling into a void
Illustration: Tradingbird

Newly unsealed court records show that OpenAI and Microsoft internally predicted their data practices would destabilize the open web, a warning they proceeded to ignore.

Recently unsealed documents from the New York Times lawsuit against OpenAI and Microsoft reveal a stark internal reality. The companies’ own employees described the massive scraping of the web as a self-destructive cycle that would harm both their models and the internet itself. These internal warnings paint a picture of organizations that understood the economic damage they were causing to publishers before proceeding with their training strategies.

The New York Times’ legal filing cites internal communications that characterize the data harvesting as an unprecedented seizure of labor. While the companies publicly framed their actions as beneficial for innovation, these private notes suggest a more calculated approach to securing the content supply chain. The stakes are high, as this dispute could define the future of how artificial intelligence interacts with human-created digital content.

Internal warnings of a destructive cycle

Microsoft’s Director of Applied Science, Brent Hecht, used strong language in internal discussions. He described the process as the largest theft of labor in history and argued that the company’s defense of fair use was a mockery of the concept. A separate internal document labeled the situation a doom loop, noting that it is unusual for a product to threaten the economic foundations of its own suppliers. This admission highlights a fundamental contradiction in the business model: relying on the very web it helps to dismantle.

Microsoft has since tried to distance itself from these specific comments. A spokesperson stated that the remarks reflected one employee’s individual perspective rather than company policy. Another executive described Hecht’s views as academic and divergent from the firm’s official stance. However, the existence of these documents in a court filing makes it difficult to dismiss them as isolated opinions, especially when they align with the broader pattern of behavior alleged by the plaintiffs.

Disavowing responsibility for known risks

The New York Times’ filing includes statements from high-profile figures like Satya Nadella and Sam Altman. Nadella acknowledged that chatbots have reduced the need for users to visit source sites directly, effectively bypassing publishers. Meanwhile, OpenAI representatives claimed ignorance regarding the removal of paywalled content from training data, despite internal admissions that the models memorized large amounts of copyrighted material. This gap between public statements and internal knowledge forms the core of the legal conflict.

According to the source material provided by GN technics/ai (en-US), the documents show that employees recognized the models were extremely good at regurgitating text verbatim. Despite this, the companies continued to expand their data collection efforts. The tension lies in the fact that the technology’s utility depends on the volume of data, while the economic viability of that data source depends on the technology not replacing it entirely.

The economic threat to publishers

The concept of a doom loop suggests a negative feedback cycle where the solution creates the problem. In this case, the AI models become more useful by consuming more web content, which in turn reduces the traffic and revenue of the websites that provide that content. As publishers struggle financially, the quality and quantity of new content may decline, potentially degrading the AI models themselves. This creates a precarious situation for the entire digital ecosystem.

The outcome of this lawsuit will likely set a precedent for how data is used in the training of large language models. If the courts find that the current practices are unsustainable, it could force a reevaluation of licensing and compensation models. For now, the industry is caught in a race between technological advancement and the preservation of the web’s economic structure.

Based on reporting by The Verge, compiled by the Tradingbird desk.

Read next

More in Tech

More from the Tech desk

All desk stories