A courtroom table displays AI copyright exhibits, a gavel, and charts showing declining website traffic.
Microsoft's AI strategy now has a paper trail, and some of the harshest criticism in it came from Microsoft's own staff. Court filings unsealed this month in The New York Times' copyright case against Microsoft and OpenAI contain internal Microsoft documents that describe the scraping behind large language models in terms usually heard from the companies' critics. Phrases like "theft" and "doom loop" appear in them.

A widely shared essay on Planet Earth and Beyond, headlined "Microsoft Knew How Horrific Their AI Push Was," reads these documents as proof that Microsoft knowingly damaged the web. The documents are real and they are striking. But the essay makes at least two factual errors and treats an unsettled legal question as closed. Below is what the record shows, what it doesn't, and why it matters to people who run Copilot, Bing and Microsoft 365.

What was unsealed, and when​

The New York Times sued OpenAI and Microsoft in the Southern District of New York in December 2023. The docket shows the complaint filed December 27, 2023 as case 1:23-cv-11195. The Times alleged that the companies built generative AI products by copying millions of its articles, and that those products now compete with it.

The new material came out on September 17, 2026. According to an analysis by Pasquale Pillitteri, federal judge Sidney H. Stein ordered the unsealing of hundreds of pages of internal documents in the lawsuit. 404 Media described the key document as a filing asking for summary judgment, in which lawyers for the New York Times laid out a series of admissions made by Microsoft and OpenAI executives in documents and depositions that until now had remained sealed or redacted.

That context matters. A summary-judgment motion asks a judge to decide issues without a trial. The characterizations in it are the plaintiffs' arguments, not findings by the court. TechCrunch also noted that much of the new information comes from the Times' own brief, not the underlying exhibits, which remain sealed. That means the quotes arrive without their original context.

Section summary: The documents are real and public. They were chosen and framed by the plaintiffs, and no court has ruled on them.

The Brent Hecht memos​

Most of the damaging quotes come from one person. In a January 2023 memo, Brent Hecht, Microsoft's director of Applied Science, described the scraping used to train AI models as "an astonishing theft of unprecedented proportions" and, possibly, "the largest theft of labor in human history". The brief quotes the fuller version: Hecht warned that "millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions."

The wording is important. The quoted sentence predicts how the public will perceive the practice. It is not a legal finding. It is still an unusually frank thing for a Microsoft science leader to write down. The memo is also dated about eleven months before the Times filed suit, so it was not written in response to the lawsuit.

The second document is dated January 2024, after the lawsuit. TechCrunch reports that an internal Microsoft presentation written by Hecht describes the decline as a "doom loop" that would "hurt the performance of our models and the entire web at the same time." The same document, as quoted in the filing, says: "It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain.'"

Pillitteri summarizes the mechanism: if Copilot pulls traffic away from news sites, those sites publish less or shut down, and future models end up with less original, quality content to learn from. Put simply, the AI business risks weakening the publishers it relies on for training material.

The Planet Earth and Beyond essay also quotes Hecht saying the fair-use defense makes "a complete mockery" of the concept. Jingletree's review of the 92-page filing reports the same thing: Hecht said Microsoft's defense makes a "complete mockery of the idea of 'fair use.'" 404 Media also reports Hecht writing that LLMs take content "without ways of distributing economic value down the supply chain, [which] necessarily threatens the economic stability of those who create the content."

The numbers behind the "doom loop"​

For Windows and Microsoft 365 readers, the most concrete part of the filing is traffic data, not rhetoric. TechCrunch reports that Microsoft's own data shows its Copilot "answer engine" caused click-through rates for The New York Times' domain to drop as much as 93% compared to traditional Bing search. TheWrap, citing the motion, gives the ranges. Click-through from Bing Chat was 87% to 93% lower for the Times' sites and 83% to 91% lower for the Daily News plaintiffs' sites, compared with traditional Bing results.

The filing also gives some sense of scale. It says OpenAI's mid-training datasets alone contain more than 91,692 copies of works published by the NYT, Daily News, and Center for Investigative Reporting, and a Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. TechCrunch reports that the brief says OpenAI gave Microsoft the entire GPT-3 training dataset. It also says Microsoft supplied training data to OpenAI through initiatives called Project Taxi and Project Mango. The Mango dataset allegedly contains at least 160,903 unique works from the plaintiff publishers.

OpenAI is under fire in the filing as well. When OpenAI researcher Nick Ryder told OpenAI president Greg Brockman about "a hack to get around nytimes paywall" to help scrape its writing, Brockman replied "ah nice." OpenAI's head of ChatGPT, Nick Turley, wrote that the company's products "are largely substitutive, period" and will become even more so as the technology improves.

Section summary: Microsoft's own measurements show that AI answers sharply cut clicks through to publishers. The plaintiffs will use that data to argue market harm, which is central to any fair-use analysis.

Where the viral essay gets it wrong​

The Planet Earth and Beyond piece makes three claims that the reporting doesn't support.

  1. The Nadella "doom loop" attribution. The essay says CEO Satya Nadella admitted Microsoft's models were creating a "doom loop." Every report of the filing attributes that phrase to Hecht's January 2024 presentation. Nadella's role is different. Testifying under oath, he described chatbots as substitutes for publisher platforms. TechCrunch quotes his deposition: conversing with chatbots "has substituted … giving you the information right there on the website on the AI platform versus needing to go to the underlying source." He also testified that "anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training." And he said that if he had known OpenAI trained on paywalled material, he would have used Microsoft's right to require OpenAI to retrain its models. That's a real concession about substitution, but it isn't a confession of theft.
  2. The wrong newspaper. At one point the essay says the New York Post is angry at Microsoft and OpenAI. The case concerns The New York Times, along with the Daily News publications and the Center for Investigative Reporting.
  3. Fair use as settled law. The essay says no sane person could consider AI training fair use. Courts haven't taken that view. TechCrunch notes that several of the new admissions run counter to OpenAI's fair use defense, particularly the requirement that a use doesn't substitute for or harm the market for the original work. The same outlet also reports that judges have so far been largely receptive to AI companies' fair-use arguments. The Department of Justice previously told a federal judge that AI training is fair use because models learn patterns from text rather than duplicate it. This is still an open fight in court.

Microsoft's response​

Microsoft doesn't accept that Hecht speaks for the company. Jingletree notes that the company has tried to distance itself from Hecht's assertions. A Microsoft spokesperson told TheWrap that the comments reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views. The spokesperson added that Microsoft's court filings explain why it considers these uses transformative and why it says Copilot doesn't substitute for publishers' journalism. On Nadella's testimony, Microsoft said he was speaking about broad changes in how people find information, and that his remarks shouldn't be confused with conclusions about the copyright questions before the court.

That's a fair point as far as it goes. A director's memo is not corporate policy. The counterpoint is also fair: Hecht leads applied science work that feeds products like Copilot, and the traffic figures came from Microsoft's own data, not from his opinion.

Why Windows and Microsoft 365 users should care​

This is my analysis, based on general industry knowledge rather than the filings.

  • Copilot's sourcing model is under pressure. Copilot in Windows, Edge and Microsoft 365 builds answers from web content. If publishers win on market harm, expect more licensing deals, more prominent citations, or both. Any of those could change how Copilot answers look and what they cost to run.
  • Referral traffic is a business problem, not just a media problem. Companies that publish documentation, knowledge bases or product content face the same loss of clicks that the Times measured. If your support site depends on Bing traffic, the 83–93% ranges above deserve attention.
  • Governance teams now have a precedent. Internal memos like Hecht's end up in discovery. Legal and compliance staff overseeing enterprise AI deployments should assume that frank internal risk assessments may one day be read aloud in court. That's a reason to act on them, not to stop writing them.

The bottom line​

"Microsoft knew" goes too far as a headline, but it isn't baseless. The record shows that at least one senior Microsoft scientist warned, both before and after the lawsuit, that the company's AI content strategy looked like uncompensated extraction and could weaken the web it depends on. It shows Microsoft's own data recording a collapse in click-throughs. And it shows Nadella agreeing under oath that chatbots substitute for publisher visits. The record doesn't show that Nadella admitted to a "doom loop," that Microsoft as a company shares Hecht's view, or that any court has found infringement.

Judge Stein still has to rule on the question that matters most: whether this is transformative use or market substitution.

 

References

  1. OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web - Jingletree jingletree.com
  2. Microsoft Knew How Horrific Their AI Push Was - planetearthandbeyond.co planetearthandbeyond.co 2026-09-27T21:01:26+00:00
  3. Microsoft exec called AI scraping ‘the largest theft of labor in human history,' new unredacted filings reveal | TechCrunch techcrunch.com