
Newsday and The Seattle Times are the latest news organizations to sue OpenAI and Microsoft for stealing their copyrighted content “without permission or compensation” and then using “the content in those datasets to train, fine-tune, and ground the large language models (‘LLMs’) underlying their generative AI products.”
The suit mirrors suits brought against OpenAI, Microsoft, and other Big Tech and Big AI companies by other new organizations, most notably The New York Times, which used OpenAI and Microsoft in December 2023.
“Like a snake eating its own tail, generative AI that is trained on painstakingly researched, expensive-to-produce content threatens to destroy the very news organizations by competing directly with them through AI-generated substitutive content,” the suit explains. “If Defendants are allowed to succeed, independent journalism of the kind Plaintiffs produce will struggle to survive.”
The suit explains that the two publications have earned a collective 30 Pulitzer Prizes while investing extensively in independent, local journalism. But in just three years, OpenAI and Microsoft have created “enormously lucrative generative in substantial part by copying, without permission or compensation, vast quantities of copyrighted journalism, including the original, expressive work of The Seattle Times and Newsday.”
The suit describes how OpenAI and Microsoft “methodically scraped” the two publications’ websites over a period spanning years using automated bots while bypassing their paywalls and “ignoring decades-old norms” to obtain their “crown jewels,” copies of their news articles. These methods are consistent with how OpenAI and Microsoft have stolen content from other publications and content creators, the suit notes, as is their practice of incorporating the stolen content into “large-scale datasets” used to improve their AI offerings. OpenAI and Microsoft also “deliberately removed and altered copyright management information (CMI) from the stolen content that includes article tiles, author names, and copyright notices.” Then, they “produce output that consists of, or contains, verbatim and paraphrased output” from the stolen works. This output “competes directly” and “acts as a substitute” for the originally produced content.
Given the specious arguments OpenAI and Microsoft have made in The New York Times suit, the Newsday/The Seattle Times suit explains that this theft is not covered by fair use because it’s not transformative in any way and is designed to compete with the publications that originally created the content. OpenAI and Microsoft “reproduce, closely paraphrase, and synthesize” the stolen content “into competing informational products that substitute for a visit to newsday.com or seattletimes.com, serving readers the same information.”
Newsday and The Seattle Times are suing OpenAI and Microsoft for violating their rights under the Copyright Act, the Digital Millennium Copyright Act (DMCA), and the Lanham Act at a federal level as well as corresponding New York and Washington state laws preventing trademark dilution. They seek to recover damages for the illegal conduct and prevent its continuation. And as with The New York Times, they have provided numerous examples of OpenAI and Microsoft regurgitating their content verbatim as proof the crimes.
In related news, The New York Times, OpenAI, and Microsoft on Friday submitted briefs in that lawsuit tied to a potential summary judgment that would eliminate the need for a time-consuming trial. Judge Sidney H. Stein will issue a ruling in the next few weeks, but it’s difficult to imagine a scenario in which OpenAI and Microsoft prevail in either of these suits.