IA · 3 September 2026 · 5 min read
US Government Sides with OpenAI: LLM Training on Protected Works Is Fair Use
In brief: The United States federal government has formally stepped into The New York Times' copyright lawsuit against OpenAI and Microsoft, filing a statement of interest that classifies large language model training as a transformative fair use. Federal attorneys argued that mandating licenses for training data would stifle scientific innovation and jeopardize the nation's technological leadership.
by Team Mocchi's
In what is widely considered the most influential legal clash shaping generative artificial intelligence, the federal government has officially taken a side. The United States administration submitted a 20-page statement of interest to a New York federal district court in defense of OpenAI and Microsoft, which were sued in December 2023 by The New York Times for allegedly scraping millions of copyrighted articles without permission or licensing fees.
As reported by TechCrunch, the federal brief frames the dispute not merely as a commercial grievance between private entities, but as an issue tied directly to national economic security. Federal attorneys highlighted that the United States holds an existential interest in maintaining global leadership in artificial intelligence, warning that a restrictive reading of copyright law would critically handicap domestic labs against foreign competitors.
Fair use and the geopolitical imperative
The core of the litigation centers on the fair use doctrine—the bedrock of US copyright law that allows uncredited and unlicensed utilization of protected works under conditions involving commentary, research, or transformative purpose. According to The Verge, government counsel cautioned that adopting the newspaper's interpretation would undermine basic copyright principles and actively hinder the constitutional objective to promote the progress of science and useful arts.
In the administration's view, large language models do not merely reproduce news reporting; instead, they digest text to learn semantic structures, conceptual hierarchies, and statistical relations, subsequently powering breakthroughs across multiple scientific and enterprise domains. Government attorneys argued this makes LLM pre-training «extraordinarily transformative», a qualification that heavily weighs in favor of fair use protection.
Between Hemingway and data scraping: the federal defense
To demystify algorithmic learning for the judiciary, the filing drew a literary parallel. As highlighted by WIRED, lawyers compared the way neural networks process vast digital corpora to a young Joan Didion retyping Ernest Hemingway’s stories at her typewriter to master sentence rhythm and storytelling mechanics. Conflating the act of reading and learning from text with copyright infringement would, according to the brief, introduce untenable legal absurdities for both software systems and human authors alike.
The publishing industry fired back immediately. New York Times spokesperson Graham James condemned the administration’s stance, stating that the White House is siding with multi-trillion-dollar corporations at the expense of creative professionals whose labor was taken without consent. James underscored that media outlets and AI ecosystems can flourish together only when AI builders fairly license the content that powers their commercial engines.
Publishers push back amid broader licensing battles
While US District Judge Sidney H. Stein is not legally bound to adopt the Department of Justice's viewpoint, legal scholars note that formal governmental intervention carries immense persuasive weight in federal courts. The outcome of this proceeding will set a definitive precedent for dozens of active lawsuits filed by news organizations, record labels, and book authors against generative AI developers.
While several prominent publishers have already negotiated private licensing pacts with companies like OpenAI, Google, and Amazon, a judicial determination validating fair use for foundation model training would fundamentally diminish content creators' bargaining power, insulating AI developers from catastrophic statutory damages and solidifying a permissive regulatory environment within the United States.
Mocchi's take
This legal development signals a turning point: the battle over AI copyright is no longer just a legal dispute, but an instrument of state-level industrial policy. If the US judiciary cements fair use for foundational training to protect domestic champions, the divergence from Europe's stricter regulatory framework—governed by the AI Act and stringent copyright directives—will widen significantly regarding model development costs. For European software teams and enterprises integrating AI workflows, the strategic takeaway is clear: while data sourcing compliance remains critical within the EU, technical differentiation will not come from building foundational models from scratch, but from mastering high-reliability software orchestration, rigorous application security, and proprietary enterprise domain data.
Further reading
- https://techcrunch.com/2026/09/02/u-s-government-sides-with-openai-on-issue-of-training-llms-on-copyrighted-material/
- https://www.theverge.com/ai-artificial-intelligence/988344/trump-administration-new-york-times-openai-lawsuit
- https://www.wired.com/story/trump-administration-sides-with-ai-giants-new-york-times-lawsuit/