The Seattle Times and Newsday have filed lawsuits against OpenAI and Microsoft, alleging that the companies used the newspapers’ journalism as training data for AI models without permission and are seeking the destruction of any copies, training datasets and models that incorporate that content.
What the complaints allege
According to The Verge, the two outlets say the company used their journalism as training data for its AI models without permission and often reproduces passages from their reporting in response to user queries. The complaint further states that the plaintiffs are asking the court to order the defendants to destroy any copies of the newspapers’ works they are holding, as well as the training datasets and AI models that incorporate them.
TechCrunch adds that the lawsuit describes generative AI as “a snake eating its own tail” that could “destroy the very organizations” that produce the content it is trained on. The filing also notes that the plaintiffs are the latest media organisations to bring copyright infringement actions against OpenAI, joining earlier suits by The New York Times, Ziff Davis, Merriam‑Webster and Encyclopedia Britannica.
Legal context and sector impact
The suits arrive at a moment when AI developers are under mounting pressure to negotiate licensing agreements for copyrighted material. OpenAI’s partnership with Microsoft, which integrates its models into the Copilot suite, has already drawn scrutiny from other publishers. If the courts grant the destruction orders sought by The Seattle Times and Newsday, the precedent could force AI firms to treat training data as a licensed asset rather than a freely harvested resource.
Industry observers note that the outcome could ripple across the broader technology sector. A ruling that compels the removal of large‑scale datasets would likely increase compliance costs for firms that rely on massive text corpora to improve model performance. It could also accelerate the emergence of formal data‑licensing marketplaces, where publishers sell access to their archives under negotiated terms.
At the same time, the lawsuits highlight a strategic tension: AI companies argue that training on publicly available text is essential for model quality, while media organisations contend that unlicensed use erodes their revenue base and threatens the viability of journalism. The “snake eating its own tail” metaphor underscores the fear that AI could render the very content creators it depends on obsolete.
Financial backdrop of the defendants
Microsoft, the larger of the two defendants, reported a net income of USD 133,749,000,000 for the fiscal year ending 30 June 2026, according to its 2026 Form 10‑K filing. The same filing shows total assets of USD 758,376,000,000 and shareholders’ equity of USD 442,387,000,000. As of the same date, Microsoft had 7,427,000,000 shares outstanding. These figures illustrate the scale of the company that could be exposed to liability if the courts enforce the destruction of AI‑trained models.
OpenAI, while privately held, is backed by Microsoft’s substantial investment and is the technology behind Microsoft’s Copilot products. The partnership means that any legal constraints on OpenAI’s data practices could directly affect Microsoft’s commercial AI offerings.
| Metric | Value | Period End | Unit |
|---|---|---|---|
| Net income | 133,749,000,000 | 2026‑06‑30 | USD |
| Total assets | 758,376,000,000 | 2026‑06‑30 | USD |
| Shareholders’ equity | 442,387,000,000 | 2026‑06‑30 | USD |
| Shares outstanding | 7,427,000,000 | 2026‑06‑30 | shares |
| Source: Microsoft 2026 Form 10‑K filing (SEC) | |||
Outlook for AI licensing and media strategy
Both lawsuits were filed in early September 2026, as reported by TechCrunch and The Verge. The timing suggests that media organisations are coordinating their legal tactics to create a unified front against AI developers. If the courts side with the plaintiffs, AI firms may need to renegotiate existing contracts and establish new licensing frameworks within months.
For the media industry, a successful injunction could provide leverage to secure compensation for past use and to set terms for future data access. However, the filings also leave several unknowns. Neither the complaints nor the research packet disclose how many specific articles or how much text the defendants allegedly used, nor do they quantify the potential impact on the AI models’ performance.
From the AI side, the companies have not publicly commented on the lawsuits beyond standard legal‑defense statements. The lack of a detailed response means the market’s reaction is currently limited to a modest dip in OpenAI‑related venture valuations, but no concrete financial impact can be measured at this stage.
Analysts anticipate that the litigation could spur a wave of settlement negotiations. Publishers may prefer licensing deals that provide ongoing revenue rather than protracted court battles. Conversely, AI developers may push back, arguing that overly restrictive licensing could slow innovation and raise costs for downstream products.
What remains unclear
The filings do not specify the exact volume of the newspapers’ content that is alleged to be embedded in the AI models, nor do they detail the technical methods used to identify “verbatim” passages. The lawsuits also do not address whether the defendants have already taken steps to delete or isolate the contested data.
Finally, the broader regulatory environment is still evolving. While the U.S. Copyright Office has issued guidance on AI‑generated works, no definitive rule exists on whether training on copyrighted text without a license constitutes infringement. The outcome of these cases could therefore shape future copyright policy.
For now, the media‑AI clash adds another layer of uncertainty to an industry already grappling with the disruptive potential of generative models. Stakeholders on both sides will be watching the courts closely, as the decision could set the terms for how AI systems access and profit from copyrighted content in the years ahead.