Legal Filing: OpenAI and Microsoft Knew They Were Stealing

Legal Filing: OpenAI and Microsoft Knew They Were Stealing

A new legal filing in the New York Times copyright infringement lawsuit against OpenAI and Microsoft introduces some troubling new evidence. Put simply, executives at both companies knew they were stealing and breaking the law, but they pushed forward regardless in a mad bid to train their AI models with high-quality content.

“This case is about, as Microsoft’s Director of Applied Science put it, ‘an astonishing theft of unprecedented proportions’ [and] perhaps the ‘largest theft of labor in human history’,” the filing opens. “[OpenAI and Microsoft] repeatedly copied millions of [The New York Times’] copyrighted articles in their entirety without permission to produce substitutive commercial AI products.”

OpenAI and Microsoft haven’t really addressed the illegality of bypassing the paywalls that The New York Times and other publishers have around their content. But the companies have publicly claimed that their theft of content from The New York Times and other content makers constitutes fair use because their AI models transform the original content into a new type of content that doesn’t directly compete with the original source. But that’s the most specious argument imaginable, and now we know that not even OpenAI and Microsoft believe it.

Among the tidbits in the filing, each of which was found during discovery, is OpenAI’s head of ChatGPT writing internally that its AI models were an “existential threat” to news publishers because they are “largely substitutive, period,” and not just transformative. Worse, they will “get more and more substitutive as they get better.” This sobering take on the content theft “eviscerates” the fair use defense, the Times writes in the filing “because substitution is ‘copyright’s bête noire’.” That is, “the goal of copyright is to provide an economic incentive to create original works.”

Lawyers representing the publication also found that OpenAI cofounder and president Greg Brockman wrote internally that its AI models were “excellent at news” and “very good at any news task” on separate occasions. And Microsoft CEO Satya Nadella wrecked havoc on the fair use defense when he noted that its AI chatbots its AI chatbots “substituted” original sources by “giving you the information right there on … the AI platform versus needing to go to the underlying source.”

“Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time,” one internal Microsoft document notes. “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain’.”

The filing goes on to document massive traffic drops at The New York Times, ZDNet, and other content makers thanks to its AI just providing the information that users want and not needing to click through to the original sources.

The New York Times is requesting a summary judgment of liability for the theft of its content, most of which is behind paywalls, the resulting AI model training that’s designed to “predict the copied articles’ expressive journalistic content,” the use of real-time AI chatbot responses built from the stolen content, and the “horse trading” by which OpenAI and Microsoft shared the content they each stole with each other.

“There is no dispute that [OpenAI and Microsoft] participated in each type of copying,” the filing adds. “And they cannot carry the burden of proving their fair use defense for any of them.”

Tagged with

Share post

Thurrott