Microsoft Exec Called AI Scraping “the Largest Theft of Labor in Human History” — Unsealed Filings Rock the NYT Lawsuit

  • AI
  • September 18, 2026

The three-year-old copyright lawsuit between The New York Times and OpenAI/Microsoft has just produced its most explosive revelation. Court filings unsealed on September 17, 2026 show that Microsoft’s Director of Applied Science, Brent Hecht, privately described the mass scraping of web content for AI training as “an astonishing theft of unprecedented proportions” — and “the largest theft of labor in human history” — in a January 2023 internal memo. Words spoken by a senior executive at the defendant itself are now shaking the foundations of the AI industry’s core legal defense.

What the Unsealed Filings Reveal

According to TechCrunch, Ars Technica and The Verge, the newly unredacted material comes from the Times’ own legal brief, while some underlying exhibits remain sealed. The filings make three explosive points: first, Microsoft and OpenAI executives privately admitted that AI training-data scraping amounted to “theft”; second, OpenAI leadership said its products posed an “existential threat” to the publishers and journalists whose work trained them; third, the companies allegedly bypassed paywalls, scraped content at massive scale, and deliberately stripped copyright notices from training data.

A pile of newspapers at a newsstand, symbolizing the battle over news content's value in the AI era
A pile of newspapers at a newsstand, symbolizing the battle over news content’s value in the AI era (AI-generated image)

The “Doom Loop”: Copilot Cut NYT Click-Through Rates by 93%

The most damning data came from Microsoft itself. Internal figures show that compared with traditional Bing search, the Copilot “answer engine” drove click-through rates for the NYT’s domain down by as much as 93%. In a January 2024 internal presentation, Hecht described the collapse of the content ecosystem as a “doom loop” that would “hurt the performance of our models and the entire web at the same time.”

Another Microsoft document stated bluntly: “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.'” A further filing acknowledged a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained.”

Nadella’s Testimony: Paywalled Content Should Be Licensed

Testifying in a deposition earlier this year, Microsoft CEO Satya Nadella said: “anything that is paywalled should be licensed by anyone who wants to use it… for grounding or training.” He made clear that had he “been made aware that OpenAI had scraped and trained on information that was behind a paywall,” he would have “invoked [Microsoft’s right to] require OpenAI to retrain its models.”

Justice statue bas-relief on the Palace of Justice as the copyright case enters a decisive phase
Justice statue bas-relief on the Palace of Justice as the copyright case enters a decisive phase. Source: Wikimedia Commons (Public domain)

91,692 Copies and Project Mango

The scale of the copying is striking. The documents reveal for the first time that OpenAI’s mid-training datasets alone contain more than 91,692 copies of works published by the NYT, the Daily News and the Center for Investigative Reporting, while a Common Crawl-derived dataset included more than 2 million documents from nytimes.com alone. A dataset assembled under a project codenamed “Mango” contains copies of at least 160,903 unique works from the news publishers.

A Paywall “Hack” and Erased Copyright Notices

The filings detail how the content was acquired: OpenAI delivered the entire GPT-3 training dataset to Microsoft, while Microsoft supplied training data to OpenAI through initiatives called Project Taxi and Project Mango. Most tellingly, when OpenAI researcher Nick Ryder told president Greg Brockman about a “hack to get around nytimes paywall,” Brockman reportedly replied: “ah nice.” Researchers also allegedly stripped copyright notices from training data before it reached the model, since they “wouldn’t want model outputting” such notices to users.

Nick Turley, OpenAI’s head of ChatGPT, wrote in internal communications that publishers face an “existential threat” from chatbot products that are “largely substitutive” and “will get more and more substitutive as they get better.” Brockman separately described the models as “excellent at news.”

Code on a computer monitor, at the heart of the dispute over AI training data
Code on a computer monitor, at the heart of the dispute over AI training data. Source: Wikimedia Commons (CC0)

Cracks in the Fair-Use Defense

Legally, US courts have so far leaned toward accepting AI companies’ “fair use” defenses, and earlier this month the Trump administration filed a brief defending OpenAI’s unlicensed use of copyrighted material. But fair use requires that the use not substitute for or harm the market for the original work — and phrases like “93% click-through decline,” “existential threat” and “largely substitutive,” coming from the defendants’ own mouths, strike directly at that requirement. Steven Lieberman, counsel for the New York Daily News, said: “The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong.” Around 160 news outlets have now joined the landmark case. Microsoft, for its part, framed Hecht as “someone on the payroll to play a contrarian,” and neither company returned requests for comment.

Conclusion: Internal Words Become the Sharpest Weapon

The outcome of this case will help decide whether the AI industry can keep training models on free web content. For creators, the unsealed filings prove that people inside the AI giants understood the problem all along. What comes next is whether courts credit these quotes — and whether the pressure finally forces the industry into genuine licensing and revenue-sharing deals. The value of content is being repriced.

Related Posts

  • September 17, 2026
Apple’s Server Comeback After 18 Years: M8 Ultra AI Servers with Nvidia NVLink Fusion Target 2029

Apple is developing an enterprise AI server built around its M8 Ultra chips in two or four-chip configurations, and has held talks with Nvidia to incorporate NVLink Fusion networking technology, The Information reports. Targeting a 2029 launch, the move would mark Apple’s first server product since Xserve was discontinued in 2011 — an 18-year absence.

  • September 16, 2026
Siri Finally Gets Smart: Apple Releases iOS 27 with Siri AI Beta, Waitlist Now Open

Apple has officially released iOS 27, bringing the long-delayed Siri AI to iPhone as a beta. With conversational answers, personal context, and on-screen awareness, the new Siri is built on Google’s Gemini models. English-only for now, waitlist and daily limits apply, the EU is left out, and users will soon be able to swap it for Claude or ChatGPT.