The thin veil of "fair use" in the AI industry has been pierced by an internal memo and a few sworn statements.
On September 17, newly unsealed filings in The New York Times copyright lawsuit against OpenAI and Microsoft put on the record what both companies have done to the news business over the past three years. Microsoft Applied Science director Brent Hecht, in a January 2023 internal memo, described OpenAI training AI on scraped news as "an astonishing theft of unprecedented proportions" and "the largest theft of labor in human history." That PowerPoint, written by Microsoft's own head of applied science, is the sharpest internal characterization of the company's own training-data practice.
Microsoft's own data answered "market substitution" first
A separate Hecht presentation logged the impact of Copilot's "answer engine" on click-through rates to nytimes.com: compared with traditional Bing search, some pages saw click-throughs collapse by as much as 93%, with the overall drop ranging 51%-94%. He called it a "doom loop" — a vicious cycle that damages both Microsoft's models and the open web at the same time.
"It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its content supply chain."
That single line is the most direct rebuttal of the "does not substitute for, or harm the market for, the original work" prong of fair use. Microsoft had the numbers in early 2024, and then spent the next three years laundering the legality of OpenAI's training corpus under the banner of fair use.
Nadella put "licensing" on the trial record
Microsoft CEO Satya Nadella, in a deposition earlier in 2025, gave another key statement: anything paywalled, if it is going to be used for grounding or for training, should be licensed; and had he known that OpenAI had scraped and trained on paywalled content, he "would have invoked [Microsoft's right to] require OpenAI to retrain its models."
That line is fatal to the Microsoft-OpenAI symbiosis. On the public stage Microsoft defends OpenAI's fair use. In court, the CEO admits that the same behavior, had he known, would have triggered a contractual demand to retrain. It is an admission that the practice is wrong, morally and contractually; it has only been allowed to continue because nobody forced the issue.
91,692 copies and "ah nice"
The technical detail is also out in the open now: OpenAI's mid-training datasets contain more than 91,692 copies of works from The New York Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset contributed more than 2 million documents from nytimes.com alone.
The black-comedy subplot is the paywall bypass. OpenAI researcher Nick Ryder messaged Brockman about "a hack to get around nytimes paywall," and Brockman replied: "ah nice." That exchange was pulled into the court filing by plaintiffs.
Why this is an industry watershed
For the past two years, courts and regulators have leaned toward fair use for AI training data. Earlier in September, the White House filed a brief siding with OpenAI. But this round of unsealed filings adds a new weight on the scale — not an outsider's critique, but the defendants' own internal documents.
The lethality of this kind of evidence is that it dismantles all three pillars of the fair-use defense at once: no market harm, non-substitutive, and non-willful. When your own director of applied science writes "doom loop," when your CEO commits that "paywalled content must be licensed," and when the partner company's president replies "ah nice" to a paywall bypass, the fair-use defense becomes very thin at both the factual and the legal layer.
What happens in the next twelve months basically depends on the settlement terms — the money, the training-data disclosure obligations, and any future "license-first" mechanism — being written into industry practice. Hecht's "largest theft of labor in human history" line may well end up being the footnote of legal record for this entire industry, rather than yet another internal gripe quietly forgotten.