Legal pressure on the leading developers of generative foundation models is mounting once again. On September 4, 2026, two historic American newspapers, The Seattle Times and Newsday, filed a joint copyright infringement lawsuit in the US District Court for the Southern District of New York. The legal complaint directly targets technology giants Microsoft and OpenAI, who are among the most influential drivers of modern artificial intelligence. With this joint legal action, two additional prominent publishing houses are challenging the widespread practice of training commercial language models on copyrighted work without consent.
This new legal challenge follows in the wake of the high-profile lawsuit previously launched by The New York Times against both companies. The entry of The Seattle Times and Newsday into the legal arena demonstrates the growing resolve of traditional media organizations to protect their intellectual property. The publishers contend that generative AI practices constitute an existential challenge to sustainable journalism, as deeply researched reporting is exploited without license agreements. Consequently, the dispute over automated data harvesting is expanding from isolated clashes into a coordinated defense across the American news industry.
At the heart of the complaint lies an allegation of extensive and calculated intellectual property theft over a multi-year timeframe. The plaintiffs assert that Microsoft and OpenAI systematically scraped hundreds of thousands of news articles without authorization to construct their proprietary systems. Crucially, the publishers allege that the defendants specifically circumvented technical barriers and digital paywalls to harvest premium reporting. According to the lawsuit, this deliberate bypassing of paywalls invalidates any claim that the content was merely indexed as part of standard, freely accessible web crawling.
According to the court filing, the harvested journalistic works were employed both to train foundation models and to power ongoing user queries. The publishers explicitly cite flagship products including OpenAI's ChatGPT and Microsoft's Copilot as direct beneficiaries of their proprietary reporting. These generative systems synthesize and reproduce reporting derived from the publishers' archives, answering user questions with information sourced directly from the newsrooms. The plaintiffs argue that this mechanism allows end users to consume their content while depriving the news organizations of web traffic and subscriber revenue.
The legal remedies demanded by the plaintiffs extend far beyond monetary compensation, introducing serious operational risks for the tech firms. Alongside financial damages, The Seattle Times and Newsday explicitly demand the court-mandated destruction of all training datasets and generative models containing their copyrighted material. Fulfilling such an order would require OpenAI and Microsoft to purge foundational weights and potentially retrain entire systems from scratch at staggering expense. A judicial mandate of this nature would strike at the technological core of modern commercial generative AI.
Legal and technology analysts view the proceedings in the Southern District of New York as a pivotal test for the future economics of generative AI development. For OpenAI and Microsoft, the lawsuit places not only historical data practices under scrutiny, but also the viability of foundation models trained on web scraping. If the court rules in favor of the publishers, it could trigger a wider domino effect of copyright claims from media organizations globally. The ultimate verdict will likely define the legal terms and licensing requirements under which artificial intelligence companies may utilize professional journalism in the future.

