← All stories
● Covered by 1 source · 1 reportMedium impact1 neutral

Microsoft Claims Minimal Copyright Infringement by Copilot in NYT Lawsuit Filings

🔄 Updated 45m ago
New to BrevFeed? We gather this story from every outlet covering it into one summary — ranked by real-world impact, not just the latest headline — so you never miss what matters. What is BrevFeed? →

Key points

  • Microsoft provided 8.2 million Copilot chat logs in discovery.
  • Fewer than 1% of logs contained 16+ words from news content.
  • Only 24 responses had 30+ matching words from books.
  • Microsoft argues this supports its fair use defense for AI training.

Microsoft's Defense Against Copyright Claims

Microsoft has filed legal documents in response to copyright infringement lawsuits brought by The New York Times and various authors. The company contends that its Copilot chatbot rarely reproduces significant portions of copyrighted content, such as news articles or books, that could substitute for the originals.

Analysis of Chat Logs

As part of the lawsuit's discovery process, Microsoft provided 8.2 million Copilot chat logs to an expert representing news publishers. These logs were specifically selected for containing keywords related to the plaintiffs' websites. Microsoft's analysis indicates that fewer than 1% of these logs, specifically 59,545, contained at least 16 words in common with news content used to train the AI model. Similarly, an expert in the authors' suit found only 24 responses with at least 30 matching words from books within the 8.2 million conversations.

The New York Times' Disagreement

The New York Times has rejected Microsoft's conclusions. Ian Crosby, lead counsel for The Times, stated that the evidence uncovered during discovery points to Microsoft and OpenAI having used The New York Times' content to create commercial products that compete with its journalism. The Times seeks accountability for what it describes as theft.

Fair Use Argument

Microsoft's argument centers on the concept of fair use, asserting that using copyrighted material for AI training datasets should fall under this legal doctrine. The company maintains that while systems like Copilot utilize copyrighted content, the resulting AI systems serve significantly different purposes than the original works. Microsoft concludes that occasional reproduction of text sections does not undermine the transformative nature of large language model training.

✨ This summary was generated by AI from the outlets' reporting listed below. It is not independently verified and may contain errors — check the original sources. How BrevFeed works →

The daily brief

One email each morning: the day's tech stories, clustered across outlets and summarized. No account needed.

One email a day. Unsubscribe in one click, any time.

Today's brief

Spend a few minutes, get the whole day. Every topic's top stories in one hands-free rundown — listen, watch, or read the transcript.

~19 min · 16 stories · Sep 04

▶ Play today's brief Listen on Spotify

New every morning, and the back catalogue is archived by date.

Reporting from

Microsoft submitted legal filings asserting that its Copilot chatbot rarely reproduces substantial portions of copyrighted material from news articles and books, citing discovery data. This claim is part of its defense against copyright infringement lawsuits from publishers, including The New York Times, and authors, arguing that AI training constitutes fair use.