OpenAI and Microsoft knew they were starting a doom loop for the web
Newly unsealed court documents in the NYT case reveal OpenAI and Microsoft executives knew AI scraping would harm publishers and the web.

Stock photo for illustration only, not from the actual event
- Unsealed court filings show OpenAI and Microsoft knew their actions would trigger a web doom loop.
- A Microsoft applied science director characterized the data scraping as the largest theft of labor in human history.
- Satya Nadella admitted that chatbots have replaced search and removed the need to visit original sources.
- Internal admissions revealed that GPT-4 memorized vast amounts of data and excels at verbatim regurgitation.
Recently unsealed court documents from the New York Times lawsuit against OpenAI and Microsoft reveal that the companies' internal documentation warned they were initiating a doom loop that would harm the web, characterized data scraping as the largest theft of labor in human history, and made a complete mockery of fair use principles.
Many of the striking quotes originated from Brent Hecht, Microsoft's Director of Applied Science. Although Microsoft attempted to distance itself from these assertions, company spokesperson Alex Haurek told The Verge that the comments reflected an individual employee's perspective and did not represent company views. Additionally, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, described Hecht's role as maintaining divergent and futuristic academic views rather than speaking for the company.
This legal battle highlights the fundamental clash between generative AI developers and original content publishers. As AI models synthesize answers directly for users without driving traffic to source websites, the traditional advertising-based revenue model for web publishers collapses, raising critical legal and ethical questions about fair use.
The 92-page filing includes numerous statements from figures such as Satya Nadella, Sam Altman, and other OpenAI employees. An internal Microsoft document explicitly stated that their AI content strategy initiated a doom loop that threatens the economic foundations of its essential content suppliers while simultaneously hurting model performance.

Stock photo for illustration only, not from the actual event
Regarding paywalled content, despite Nadella stating that paywalled material should be licensed, an OpenAI representative admitted ignorance of any efforts to detect or remove paywalled content from training datasets. Furthermore, while acknowledging the importance of preventing copyright violations, employees admitted that GPT-4 memorized massive amounts of data and excelled at regurgitating copyrighted text verbatim.
"Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time"
Microsoft Internal Document
Source: The Verge
Found something wrong in this article? Report an issue with this article
Comments
Leave a Comment