NYT Case: A Microsoft Exec Called Scraping Outright Theft
Filings unsealed Thursday in the New York Times case show a Microsoft executive internally describing scraping-based training as the largest theft of labour in human history. Those words land squarely on what a fair use defence has to prove.

The most closely watched court case in generative AI just lost a chunk of its black ink. On Thursday 17 September 2026, the summary judgment motion filed by the news organisations led by the New York Times was unsealed, in the three-year-old suit against OpenAI and Microsoft. Ars Technica and TechCrunch both reach the same conclusion: what the two companies argue in front of a judge bears little resemblance to what they wrote to each other.
A January 2023 memo, and the word you never put in writing
Brent Hecht, Director of Applied Science at Microsoft, described the mass harvesting of news content, in an internal memo from January 2023, as “an astonishing theft of unprecedented proportions”, and possibly “the largest theft of labor in human history”. That is the word theft, typed into a corporate document, by a director, on purpose. In another text, he judged that the plan to scrape the press at scale makes “a complete mockery of the idea of ‘fair use’”.
The same document concedes that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use”. According to Ars Technica, Microsoft now distances itself from its executive’s remarks.
The “doom loop” shows up in the click-through numbers
An internal presentation from January 2024, by the same Hecht, puts a number on the damage: the click-through rate to the New York Times domain drops by as much as 93% with the Copilot answer engine, compared with a standard Bing search. Ars Technica notes falls of 83 to 93% for some plaintiffs, 51 to 94% for others. The document describes a “doom loop” that will “hurt the performance of our models and the entire web at the same time”, then states the problem with a clarity rarely seen on an internal slide: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain’.”
Over at OpenAI, Nick Turley, who runs ChatGPT, writes that publishers face an “existential threat”, that the products are “largely substitutive, period” and “will get more and more substitutive as they get better”. An in-house engineer adds: “no matter how prominently we show the links, users won’t click.”
Why these sentences hit exactly where it hurts
Fair use permits the use of a protected work in certain cases — parody, news reporting, criticism — provided, among other things, that the use does not substitute for the market of the original. Which is precisely the vocabulary used internally: substitutive, existential threat, traffic collapse. Satya Nadella also testified under oath earlier this year that “anything that is paywalled should be licensed by anyone who wants to use it…for grounding or training”, and said that had he known OpenAI had trained on paid content, he would have demanded the models be retrained.
The scale fits in three numbers: over 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting in OpenAI’s mid-training sets alone; over 2 million documents from nytimes.com in a corpus derived from Common Crawl; at least 160,903 unique works in the dataset assembled through the Mango project.
What is still sealed
One caveat worth its weight, flagged by TechCrunch: most of this comes from the plaintiffs’ brief, not from the exhibits themselves, which remain sealed. The quotes are circulating stripped of their original context, and no internal memo has ever won a lawsuit on its own. So far, TechCrunch recalls, judges have tended to be receptive to the fair use argument, and the Trump administration filed a brief this month defending OpenAI’s unlicensed use of protected content.
Three years of litigation, and the smoking gun turns out to have been sitting in a PowerPoint deck the whole time.
Sources (2)
Written with AI assistance from the sources cited above, then reviewed and approved before publication by Sébastien Soulier.


