The New York Times' copyright lawsuit against OpenAI and Microsoft has, for the first time in three years, exposed internal statements of this gravity. On September 17, TechCrunch obtained newly unsealed filings showing that Microsoft's director of Applied Science, Brent Hecht, called AI data scraping 'an astonishing theft of unprecedented proportions' and 'the largest theft of labor in human history' in an internal memo dated January 2023. The same filing also disclosed concrete training-set scale and commercial impact — details that until now had only lived in speculation.
What the filings reveal: training-set scale and commercial harm, both now quantified
What this unsealing does is turn two long-disputed claims from copyright litigation into numeric accusations:
- Training data scale: OpenAI's mid-training datasets contain more than 91,692 copies of works published by The New York Times, the Daily News, and the Center for Investigative Reporting. A Common Crawl-derived dataset alone pulled more than 2 million documents from nytimes.com.
- Copilot's self-harm effect: Microsoft's own data shows its Copilot 'answer engine' caused click-through rates on The New York Times' domain to drop as much as 93% compared with traditional Bing search. A separate January 2024 internal Microsoft presentation by Hecht called this a 'doom loop' that would 'hurt the performance of our models and the entire web at the same time.'
- Bidirectional data flow: Codenamed projects like Project Taxi and Project Mango moved training material between OpenAI and Microsoft; the Project Mango dataset contained at least 160,903 unique works from the three plaintiff publishers.
- Paywall circumvention: When OpenAI researcher Nick Ryder mentioned 'a hack to get around nytimes paywall' to president Greg Brockman, Brockman reportedly replied: 'ah nice.'
These unsealed materials were not disclosed by OpenAI or Microsoft themselves — they were quoted by The New York Times in its own brief; the underlying exhibits remain sealed. In other words, the quotes are presented out of original context, and the companies still retain room to push back.
Personal take: what is genuinely dangerous about this case is not 'where the data came from'
Technical readers tend to get caught up in the 'how much did they actually scrape' wow factor, but what actually moves the industry is that it has directly dismantled the two most critical pillars of the fair-use defense. Fair-use jurisprudence requires AI companies to show that 'the use does not substitute for or harm the market of the original work' — but the unsealed filings show Microsoft itself citing 83%-93% click-through declines, and OpenAI internally describing ChatGPT as 'largely substitutive,' actively admitting substitution effects. Add to that Hecht's literal admission that 'an end-product threatens the economic foundations of its essential suppliers,' and the 'no substitution' prong of the fair-use defense has collapsed at the factual level.
For teams currently working on model training and productization, what is worth tracking is not the verdict itself, but three things: first, whether courts will incorporate 'AI companies' internal self-knowledge of market harm' into fair-use analysis; second, whether subsequent settlement terms will require transparency disclosure of training data; third, whether codename-style data flows like Project Mango will be retroactively classified as 'collusive infringement.' Regardless of how the final ruling comes down, this batch of filings has already been added to the evidence list for follow-on litigation by other news organizations — the middle innings of the copyright war have turned a page.
So what
If you are AI company legal or compliance staff, you should now put this scenario into internal drills: do your training-data source documents contain emotive phrasing like 'we know this is theft'? If you are on a data-acquisition or dataset team, Project Mango's naming convention and the researcher narration of 'paywall circumvention' should become the new compliance red lines — they are evidence points for 'knowing' intent. If you are merely a reader following the AI industry, what is worth tracking in the months ahead is not a particular damages number, but whether the boundary of fair use will be redrawn in 2027 by this batch of filings.
Source: TechCrunch reporting by Rebecca Bellan, September 17, 2026.