IA · 6 September 2026 · 5 min read

The AI Copyright War Escalates: New Publishers Sue OpenAI as Microsoft Reveals Copilot Chat Logs

In brief: Two prominent regional US newspapers, The Seattle Times and Newsday, have filed a joint federal lawsuit against OpenAI and Microsoft over unauthorized copyrighted content used for AI training. Meanwhile, fighting ongoing claims from The New York Times, Microsoft submitted forensic data from over eight million Copilot conversations to argue that chatbot queries rarely reproduce copyrighted prose, backing its fair-use defense.

by Team Mocchi's

The AI Copyright War Escalates: New Publishers Sue OpenAI as Microsoft Reveals Copilot Chat Logs

The legal confrontation between legacy publishers and artificial intelligence developers is entering an unprecedented phase. Two major regional American newsrooms, The Seattle Times and Newsday, have filed a joint federal lawsuit against OpenAI and Microsoft, alleging that decades of investigative reporting were appropriated without permission or compensation to build frontier generative models.

As reported by TechCrunch, the legal complaint characterizes the current trajectory of generative AI as «a snake eating its own tail», warned to devour the very reporting organizations whose intellectual output sustains it. The publishers contend that systems such as ChatGPT and Microsoft Copilot act as rapacious consumers rather than helpful research tools, serving derivative summaries that siphon web traffic and commercial advertising revenues away from newsrooms.

Microsoft’s empirical defense: auditing 8.2 million Copilot logs

While fresh litigation lands in court, the consolidated proceedings spearheaded by The New York Times, the Authors Guild, and the Center for Investigative Reporting (CIR) have reached an evidentiary milestone. Seeking a summary judgment, Microsoft submitted empirical user data in federal court to counter the claim that conversational assistants displace original journalism.

According to The Verge, Microsoft disclosed a targeted dataset of 8.2 million Copilot chat transcripts to an independent expert representing the plaintiffs. The records were filtered by keywords closely tied to the publishers' domains. Even within this targeted pool, the audit revealed that only 59,545 conversations—fewer than one percent—contained sequences of at least 16 words matching the underlying news reporting.

The findings were even starker regarding book authors: an expert reviewing the same dataset identified merely 24 assistant responses that matched 30 or more consecutive words, touching only 10 out of the 212 published books evaluated. Microsoft's defense team argues these figures demonstrate that Copilot serves a transformative purpose rather than acting as a market substitute, squarely placing AI training and reference retrieval within the doctrine of fair use.

Shifting alliances and the fair use dilemma

The lead counsel for The New York Times, Ian Crosby, pushed back against Microsoft's conclusions, stating that discovery documents confirm the unauthorized ingestion of original journalism to fuel commercial products. The plaintiffs maintain that statistical infrequency of regurgitation does not eliminate liability for wholesale pre-training ingestion.

The inclusion of The Seattle Times highlights how quickly industry relations have deteriorated. Both Microsoft and OpenAI had previously sponsored journalism grants and technical fellowships with the Seattle-based publication. A Microsoft spokesperson expressed surprise over the lawsuit, emphasizing willingness to negotiate, but the lawsuit makes clear that philanthropic grants are no longer enough to replace formal, high-value licensing frameworks.

Mocchi's take

This expanding legal battlefield signals a turning point for any company deploying generative tools and custom AI workflows: data compliance is shifting from abstract ethical debate to forensic operational reality. When architecting internal retrieval systems or external bots, engineering teams must maintain strict provenance over source documents and verify licensing boundaries to prevent unintended verbatim outputs. As Microsoft's legal disclosure illustrates, statistical logging and auditability will soon become mandatory standards to demonstrate legitimate AI usage.

Further reading

All articles on the Mocchi's blog