AI logos and data streams connect news archives to a courthouse, legal papers, and a judge’s gavel.
The Shoestring has become a named plaintiff in a new, separate federal copyright case against Microsoft and OpenAI, filed September 16 in the Southern District of New York. The suit is not merely an addition to the earlier local-newspaper action filed in June: the docket identifies a new case, Times Publishing Company et al. v. Microsoft Corporation et al., while also asking the court to treat it as related to the broader OpenAI news-publisher multidistrict litigation.

The Shoestring reported that it and other independent publishers sued over the alleged uncompensated use of their reporting in ChatGPT and Microsoft Copilot. The court record confirms that its apparent corporate operator, Noisy Creek Inc., is among 26 plaintiffs in the September complaint. It also confirms Microsoft is sued alongside 10 OpenAI-related entities, including OpenAI Inc., OpenAI LP, OpenAI Group PBC, and the OpenAI Foundation.

For Windows and Microsoft customers, the case matters less because it threatens an immediate Copilot change than because it sharpens the legal question underlying the company’s generative-AI strategy: whether building and operating commercial AI systems on copied web material is protected fair use, or whether training and output behavior create compensable copyright liability.

A new filing, not an amendment to June’s newspaper case​

The Shoestring’s account places its suit alongside the New York Times litigation and the earlier case brought by publishers of hundreds of local newspapers. That broad framing is accurate, but the procedural detail is more important than it first appears.

The September 16 filing has its own case number, 1:26-cv-08082, and its own complaint. It is led by Times Publishing Company, the parent of the Tampa Bay Times, and includes the Austin Chronicle, Cape Gazette, SwimSwam, the Boston Institute for Nonprofit Journalism’s Massachusetts Media Fund, and Noisy Creek. The docket shows 26 legal plaintiffs, rather than a single class action in which every local publisher automatically participates.

Platkin LLP, founded by former New Jersey Attorney General Matthew Platkin, represents the plaintiffs. The firm previously filed the June case, Richner Communications Inc. v. Microsoft Corp., which Bloomberg Law reported involved publishers operating nearly 400 newspapers. Subsequent reporting from Insider NJ and the Boston Institute for Nonprofit Journalism describes the two Platkin-led groups together as representing more than 550 publications.

Those numbers should not be confused. More than 550 is a description of the outlets connected to the combined publisher campaign; the September filing itself is a complaint by 26 corporate or nonprofit plaintiffs. It also is not a judicial finding that every one of those outlets’ articles was used in an OpenAI model or in Microsoft Copilot.

The new plaintiffs filed a statement asking for the case to be deemed related to the existing news-publisher multidistrict litigation, designated 25-md-3143. That makes consolidation or coordinated pretrial handling likely, but it has not yet turned the September filing into the June lawsuit. The distinction affects what evidence, damages claims, and registered copyrights are actually before the court.


The complaint’s token-count evidence is an allegation, not a model audit​

The Shoestring says the evidence includes more than 68,000 tokens of its text in the dataset at issue. The September court docket confirms that the plaintiffs filed an exhibit titled “C4 Token Count by Domain,” meaning the claim is tied to a measurable corpus rather than only to a general accusation that AI companies scraped the web.

But token counts need careful reading. A token is a unit of text processed by a language model; it is not a count of articles, page views, model outputs, or damages. And the presence of a publisher’s material in C4 — the filtered English-language corpus derived from Common Crawl web data — does not by itself establish that a particular OpenAI model trained on every listed token, that Microsoft independently copied it, or that a user can reproduce it through Copilot.

That gap is central to the case. The plaintiffs need to connect alleged copying to protected works, show a legally actionable role for Microsoft, and overcome the defendants’ expected fair-use argument. A public corpus may help demonstrate that material was available for training or that a web crawl captured a domain, but it is weaker than an authenticated record of a particular training run or a reproducible output that delivers protected text.

One account of the filed complaint, published by Agenccy AI after reviewing the court document, says the exhibit lists more than 38 million C4 tokens across the 26 plaintiffs’ domains and assigns Noisy Creek a substantially higher figure than the 68,000 quoted in The Shoestring’s own report. That discrepancy may reflect different datasets, collection periods, domains, or counting methods. The public docket identifies the token-count exhibit but does not display its contents, so there is not yet enough verified public evidence to say which number applies to The Shoestring’s articles specifically.

The useful takeaway is not that a token total proves infringement. It is that these plaintiffs are attempting to make the training-data dispute concrete, domain by domain, rather than leaving it at the level of an industry-wide complaint.

Copyright claims are paired with a DMCA allegation​

The complaint asserts direct copyright infringement, vicarious copyright infringement, and a claim under the Digital Millennium Copyright Act. The docket identifies an additional exhibit labeled “CMI Stripping Examples,” referring to copyright-management information such as author names, copyright notices, and other identifying metadata.

The DMCA count may be more consequential than it sounds. Copyright infringement claims in AI cases frequently collide with a difficult fair-use debate: whether training is sufficiently transformative, whether the source material was publicly accessible, how much was copied, and whether AI answers substitute for the original work. A claim that copyright-management information was deliberately removed has a different statutory basis and may be less easily resolved by broad arguments about transformative training.

Still, it remains an allegation. The plaintiffs must show the required state of mind and a causal connection between removing that information and enabling or concealing infringement. A model’s answer lacking an author byline is not automatically proof that a defendant removed metadata from a specific article in a legally actionable way.

The complaint seeks statutory damages, compensatory damages, restitution, disgorgement of profits, injunctive relief, and attorney fees. It does not state a total dollar demand in the public docket description. Assertions that the case seeks a defined multibillion-dollar payout would therefore go beyond the filed record.

There is also a narrower wrinkle in the requested remedies. Agenccy AI’s review of the complaint says the explicit request to remove registered works from models and training datasets is framed around Times Publishing Company’s registered works, not a blanket injunction requiring the removal of every article from all 26 plaintiffs. If accurate, that underscores a practical challenge for smaller publishers: copyright registration status can determine which remedies are available and how broadly a court can order removal or destruction.


Microsoft’s exposure runs through Copilot and its OpenAI partnership​

The new complaint treats Microsoft as more than a distant investor. It names Microsoft as a defendant and alleges that the disputed material helped build and commercialize products including Microsoft Copilot as well as ChatGPT.

That makes the litigation relevant across Microsoft’s consumer and enterprise AI portfolio. Copilot is embedded in Windows, Microsoft 365, GitHub, Azure, Bing, and security and administration workflows. A ruling that places licensing obligations or damages exposure on a company deploying models could change the economics of which models Microsoft can offer, the data provenance documentation customers demand, and the contractual protections required in enterprise AI agreements.

It would not mean that Copilot suddenly stops functioning or that organizations should expect existing subscriptions to be withdrawn. This case is in its opening stage. There has been no merits ruling, no injunction, and no court order requiring Microsoft to alter Copilot, Azure OpenAI Service, or Microsoft 365 Copilot.

OpenAI’s public position, reported by Bloomberg Law in connection with the June publisher case, is that its models are trained on publicly available data and are grounded in fair use. Microsoft did not provide a response to Bloomberg Law at that time. The new docket, filed only September 16, contains the complaint and attorney appearances but no answer from Microsoft or OpenAI.

For IT decision-makers, the immediate issue is contractual rather than operational: vendors’ assurances that an AI service indemnifies customers for output-related copyright claims do not resolve claims over the data used to build the underlying model. Those are separate risks, with different parties and remedies.

The government has entered the larger fight on OpenAI’s side​

The September filing arrives as the New York Times case against Microsoft and OpenAI approaches a more consequential stage. The Times sued in late 2023, and its case survived a major motion-to-dismiss challenge. Axios reported this month that the parties have now submitted summary-judgment arguments to Judge Sidney H. Stein, who could decide whether the matter proceeds to trial.

On September 2, The New York Times reported that the Justice Department filed a statement of interest supporting OpenAI’s position. The department argued that training on copyrighted material can be transformative fair use and said the benefits of AI outweigh competitive harm to publishers. The department linked its position to national security and U.S. AI leadership.

That intervention does not decide the case, and it does not erase publishers’ rights. It does signal that the federal government is pressing a policy argument that directly favors OpenAI and Microsoft’s defense: broad liability for model training, the government contends, could damage American AI development.

The September publishers have therefore entered a legal fight where the defendants have not only substantial technical and financial resources, but also support from the Justice Department on the central fair-use question. Their stronger factual claims will need to be specific: particular copying, particular works, particular output behavior, and evidence of how Microsoft’s products benefited from the alleged conduct.

The first concrete milestone is procedural. The Southern District of New York must determine how the Times Publishing case will be handled alongside the existing OpenAI news litigation. Until then, Noisy Creek and its fellow plaintiffs have put Microsoft and OpenAI on notice — but a token-count exhibit and a 52-page complaint are the beginning of proof, not the end of it.