Unsealed court filings in the New York Times lawsuit reveal that internal OpenAI and Microsoft documents from 2023‑2024 explicitly warned that their large‑scale web‑scraping for large‑language‑model training had created a “doom loop” that would simultaneously reduce traffic to publishers and undermine the future quality of the models themselves.
The unsealed memos
The Verge reported that an internal Microsoft document stated:
“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time.”The same set of filings described the practice as “the largest theft of labor in human history.”
Microsoft’s Director of Applied Science Brent Hecht is quoted in the documents calling the harvesting a “largest theft of labor in human history” and warning that the AI products “make a complete mockery of the idea of fair use.”
OpenAI’s head of ChatGPT, identified in the filings as Nick Turley, is cited as observing that once a user receives an answer, there is “no good reason to click” the source link, underscoring the reduced incentive for users to visit original publisher sites.
Company context
OpenAI, founded in December 2015 and headquartered in San Francisco, is led by CEO Sam Altman and employs roughly 4,500 staff, according to Wikidata. Microsoft, incorporated in 1975 and based in Redmond, Washington, is led by CEO Satya Nadella and reports about 221,000 employees. Both firms are among the largest AI developers in the United States.
Financial backdrop
While the filings focus on strategic risk, Microsoft’s most recent public financial statements provide a snapshot of the scale at which the companies operate. The table below pulls three key metrics from Microsoft’s 10‑K filing for the fiscal year ended 30 June 2026 and a 10‑Q filing for the quarter ended 31 December 2010.
| Metric | Value | Period end | Unit | Source |
|---|---|---|---|---|
| Revenue | 36,148,000,000 | 31 Dec 2010 | USD | Microsoft 10‑Q (filed 27 Jan 2011) |
| Net income | 133,749,000,000 | 30 Jun 2026 | USD | Microsoft 10‑K (filed 29 Jul 2026) |
| Total assets | 758,376,000,000 | 30 Jun 2026 | USD | Microsoft 10‑K (filed 29 Jul 2026) |
These figures illustrate the massive scale of Microsoft’s balance sheet and earnings, underscoring why the company’s data‑harvesting strategy matters to the broader web ecosystem.
Implications for publishers and AI models
The memos suggest a two‑fold risk. First, by scraping large swaths of publicly available content, the AI systems reduce the number of clicks that would otherwise drive ad revenue to news sites and other publishers. Brent Hecht’s comment that the practice is “the largest theft of labor in human history” frames the issue as a systematic extraction of value without compensation.
Second, the same documents warn that the feedback loop could degrade model performance. If the scraped data become stale or if publishers lock down their content, future training sets may be poorer, leading to “hurt the performance of our models,” as the Microsoft memo puts it. OpenAI’s observation about users lacking incentive to click source links reinforces this dynamic.
Both firms have not publicly responded to the filings, and the court documents do not quantify the magnitude of traffic loss or model degradation. The risk remains largely qualitative, based on internal assessments rather than external measurement.
What remains unknown
- The exact volume of traffic diverted from publishers to AI‑generated answers.
- Quantitative estimates of how model accuracy or relevance will decline if the “doom loop” persists.
- Whether either company has altered its data‑collection practices since the internal warnings.
Regulators and industry observers will likely seek more concrete data as the litigation proceeds. For now, the unsealed documents provide the first public proof that senior AI leaders were aware of the self‑reinforcing harms and nevertheless continued the scraping strategy.
Next steps
The New York Times case is still pending, and the court may order further discovery that could shed light on any remedial actions taken by OpenAI or Microsoft. Publishers are watching closely, as any shift in AI data‑use policy could affect the economics of online news. Meanwhile, the AI community faces a policy dilemma: balancing the need for large, diverse training data against the sustainability of the web ecosystem that supplies it.