IA · 18 September 2026 · 5 min read

‘The Largest Theft of Labor in History’: Unsealed Docs Expose Microsoft and OpenAI’s Internal Fears

In brief: Unsealed records from the copyright lawsuit filed by The New York Times have exposed internal communications between Microsoft and OpenAI leadership. The documents reveal researchers describing news scraping as an unprecedented theft, warning of a 'doom loop' destroying the web's content supply chain, and reporting referral traffic collapses of up to 94%. These revelations directly threaten the companies' primary fair use legal defense.

by Team Mocchi's

‘The Largest Theft of Labor in History’: Unsealed Docs Expose Microsoft and OpenAI’s Internal Fears

In courtrooms and public statements, tech conglomerates developing generative artificial intelligence have consistently maintained that training large language models on billions of public web pages constitutes legitimate fair use under US copyright law. Behind closed doors, however, their internal assessment was drastically different. That reality has now been laid bare by a massive trove of unsealed internal documents in the federal lawsuit pitting The New York Times and fellow publishers against Microsoft and OpenAI.

As reported by Ars Technica, newly unredacted filings supporting a motion for summary judgment expose candid email exchanges and technical memos previously shielded from the public. In these records, key researchers and product heads from both companies characterized the automated harvesting of news articles not as fair technological progress, but as an unsustainable form of commercial appropriation.

Internal Admissions: From 'Astonishing Theft' to Mocking Fair Use

Among the most damaging evidence are internal memoranda authored by Brent Hecht, Microsoft's Director of Applied Science. In his writings, Hecht repeatedly warned colleagues that scraping news content to train generative models was «an astonishing theft of unprecedented proportions», stating it was perhaps «the largest theft of labor in human history».

Hecht went on to directly dismantle the public argument his company would later present to judges, noting that the broad news-scraping strategy made «a complete mockery of the idea of 'fair use'». He acknowledged that «almost no one intended for content they created to be used in this fashion, nor are they compensated for its use».

Over at OpenAI, internal sentiment echoed identical concerns. Nick Turley, head of product for ChatGPT, admitted in internal communications that news publishers were facing an «existential threat». The danger stemmed directly from the fact that commercial chatbots trained on journalistic reporting were operating as outright market substitutes for the underlying news providers.

The 'Doom Loop' and Evaporating Referral Traffic

One of the most consequential insights from the unsealed filings is the formal discussion of an AI «doom loop». A Microsoft internal document explained the structural paradox: «It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'».

Engineers and strategists understood that by choking off the revenue streams sustaining professional journalism, generative tools risked destroying the very pipeline of accurate, fresh data required to keep their models functional, degrading the quality of the broader open web.

Empirical data gathered internally by Microsoft and OpenAI confirmed these fears. The filings completely undermine the narrative that chatbots drive valuable discovery traffic back to publishers. For several news organizations suing the firms, click-through rates from search-assisted conversational agents plunged between 83% and 93%, while others recorded drops ranging from 51% to 94%. Chatbots do not direct readers to the original articles; they synthesize and quote verbatim excerpts, rendering visits to publisher sites superfluous.

A Turning Point in the Copyright Reckoning

The unsealing of these records fundamentally shifts the legal battleground. Under US fair use doctrine, the fourth statutory factor assesses the effect of the unauthorized use on the potential market for or value of the copyrighted work. Written confessions proving that tech executives knowingly built products designed to be «largely substitutive» severely undermine their defensive claims.

Publishers are now approaching trial with an unprecedented evidentiary record: conclusive proof that generative AI leaders understood the economic harm they were inflicting and recognized that their legal defenses contradicted their own internal assessments.

Mocchi's Take

For businesses and engineering teams building software or adopting AI systems, these unsealed documents mark the definitive end of the 'free scraping' era. When the architects of frontier models privately concede that indiscriminate training was legally and economically untenable, relying on black-box tools with dubious data provenance introduces severe compliance vulnerabilities. For organizations holding proprietary technical knowledge, industrial documentation, or specialized databases, this shift creates unprecedented leverage to demand fair licensing agreements rather than giving away value for free. As a software agency, we believe enterprise AI must move away from predatory scraping and embrace auditable architectures: governed RAG pipelines anchored in verified proprietary data, coupled with technology providers that offer clear legal indemnities.

Further reading

All articles on the Mocchi's blog