Microsoft claims Copilot rarely reproduces content from NYT articles

Microsoft official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Microsoft claims Copilot rarely reproduces content from NYT articles

The tech giant is wielding 8.2 million chat logs as evidence in its high-stakes copyright battle with news publishers.

Microsoft wants a federal court to know that its AI chatbot is not a plagiarism machine. In new legal filings responding to copyright claims from The New York Times and other publishers, the company argues that Copilot rarely reproduces even full sentences from news articles, let alone chunks substantial enough to substitute for the original work.

The argument rests on a massive trove of data: roughly 8.2 million Copilot chat logs that Microsoft handed over to an expert hired by the news publishers during discovery. And the company says these logs were deliberately selected to be the most damning ones possible.

Advertisement

The numbers behind the defense

Microsoft’s legal strategy here is essentially statistical. The company claims the 8.2 million logs were filtered specifically because they contained keywords implicating use of the news plaintiffs’ websites. In other words, these weren’t randomly pulled conversations about weekend dinner recipes. They were the outputs most likely to contain reproduced content from the publishers suing Microsoft.

Out of that filtered pool, Microsoft says the analysis identified 59,545 instances where Copilot’s output overlapped with plaintiffs’ content. That’s roughly 0.7% of the total sample, a number Microsoft clearly considers exculpatory.

A lawsuit with broad implications

The underlying case, consolidated under NYT v. Microsoft et al. in the US District Court for the Southern District of New York, was originally filed in December 2023. The Times alleged that Microsoft and OpenAI used its journalism without permission to train the AI models powering Copilot and ChatGPT, effectively creating tools that could compete with the original reporting.

The case has evolved significantly since then. Discovery disputes intensified through late 2025 and into early 2026, with battles over how chat logs should be produced and what formats would be accessible to the plaintiffs’ experts. In June 2026, The New York Times amended its complaint to place even greater emphasis on Microsoft’s specific role in facilitating the use of copyrighted content, rather than treating the company as merely OpenAI’s business partner.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.
Microsoft claims Copilot rarely reproduces content from NYT articles
Microsoft claims Copilot rarely reproduces content from NYT articles

The tech giant is wielding 8.2 million chat logs as evidence in its high-stakes copyright battle with news publishers.

Microsoft official logo (public domain, Wikimedia Commons) — CryptoBriefing brand treatment

Microsoft wants a federal court to know that its AI chatbot is not a plagiarism machine. In new legal filings responding to copyright claims from The New York Times and other publishers, the company argues that Copilot rarely reproduces even full sentences from news articles, let alone chunks substantial enough to substitute for the original work.

The argument rests on a massive trove of data: roughly 8.2 million Copilot chat logs that Microsoft handed over to an expert hired by the news publishers during discovery. And the company says these logs were deliberately selected to be the most damning ones possible.

Advertisement

The numbers behind the defense

Microsoft’s legal strategy here is essentially statistical. The company claims the 8.2 million logs were filtered specifically because they contained keywords implicating use of the news plaintiffs’ websites. In other words, these weren’t randomly pulled conversations about weekend dinner recipes. They were the outputs most likely to contain reproduced content from the publishers suing Microsoft.

Out of that filtered pool, Microsoft says the analysis identified 59,545 instances where Copilot’s output overlapped with plaintiffs’ content. That’s roughly 0.7% of the total sample, a number Microsoft clearly considers exculpatory.

A lawsuit with broad implications

The underlying case, consolidated under NYT v. Microsoft et al. in the US District Court for the Southern District of New York, was originally filed in December 2023. The Times alleged that Microsoft and OpenAI used its journalism without permission to train the AI models powering Copilot and ChatGPT, effectively creating tools that could compete with the original reporting.

The case has evolved significantly since then. Discovery disputes intensified through late 2025 and into early 2026, with battles over how chat logs should be produced and what formats would be accessible to the plaintiffs’ experts. In June 2026, The New York Times amended its complaint to place even greater emphasis on Microsoft’s specific role in facilitating the use of copyrighted content, rather than treating the company as merely OpenAI’s business partner.

Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.