Tech circles were rocked this week by an unsealed court filing. In the New York Times vs OpenAI and Microsoft copyright lawsuit, the plaintiff's legal team submitted a 92-page unredacted summary judgment motion packed with quotes from internal documents and sworn testimony by Microsoft and OpenAI executives — statements that Microsoft and OpenAI had kept sealed for over three years under "trade secret" claims.
The most explosive quote came from Microsoft Director of Applied Science Brent Hecht. In an internal memo from January 2023, he wrote that scraping news content to train AI was "the largest theft of labor in human history", calling it "a complete mockery of the fair use doctrine." Elsewhere, he described LLM training as "stealing content without distributing economic value down the supply chain, which necessarily threatens the economic stability of those who create the content."
What the documents say
Even more damning is a Microsoft-authored strategy document that admits the AI business has entered a "doom loop". The document reads: "Our AI content strategy has started a 'doom loop' that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its 'content supply chain'." That same strategy document also concludes: "LLMs are a product that destroys its own supply chain."
Microsoft's own data backs this up. The filing states that click-throughs to some publisher sites fell by 83-93%, and to other news organizations by 51-94%. In sworn testimony, Microsoft CEO Satya Nadella admitted that after scraping the New York Times and other news sites, clicks to those sites on Bing "completely cratered", dropping over 90%. In other words, AI used their content to train better search engines — and those search engines then squeezed the same content providers out of search results. Nadella also confirmed: chatbots have essentially "substituted" the act of going to original websites for information.
OpenAI's side of the story is no better. The filing reveals that when co-founder Greg Brockman was told their crawler had bypassed the New York Times paywall, he replied "Ah, nice". An OpenAI corporate representative testified he was "unaware" of any effort by the company to "detect paywalled content" in training data. Head of ChatGPT Nick Turley wrote internally that AI chatbots pose an "existential threat" to publishers because they "are largely substitutive". OpenAI policy director Jack Clark added: "We are creating systems that substitute for the labor of the people that define the 'culture' of society."
Why this filing matters
The filing's significance is not that it reveals facts the world didn't already know. What matters is that these admissions come from the defendants' own mouths. Microsoft and OpenAI's core legal defense has been "fair use" and "transformative use" — that training an LLM is highly transformative and does not substitute for the original work. But their own executives, in internal documents on the same timeline, described these models as "theft", "labor substitution", "doom loop", and "supply chain destruction". That gap between internal admission and external defense is exactly the ammunition the NYT legal team is using to argue for summary judgment.
Industry impact and the deeper question
The filing's impact extends far beyond this single lawsuit — it is the first time AI companies' internal assessments of their own training data compliance have been put on the public record at this scale. In the coming months, debates around training data licensing, publisher "opt-out" mechanisms, and whether AI-powered retrieval should share revenue with content sources will accelerate. OpenAI's own economic expert also admitted that Google's introduction of AI Overviews may have depressed search referrals to publishers by 20 to 60 percent, suggesting that the entire AI retrieval ecosystem's squeeze on content industries is systemic.
An even deeper signal: even if AI companies win the "fair use" argument, the "doom loop" problem doesn't disappear. LLM quality depends on high-quality original content — if AI destroys the content industry's business model, where does the next generation of training data come from? That is a question the AI industry cannot leave solely to its lawyers.