"Theft of labor" leaked from a Microsoft internal PPT
Ars Technica recently obtained a Microsoft internal document whose main author is Brent Hecht, Microsoft's Director of Applied Science. His wording was unusually blunt: scraping news content to train AI is "the largest theft of labor in human history" — he used the phrase "theft of labor" in the file and explicitly said this "completely mocks the concept of fair use." When one of the world's most valuable tech companies acknowledges, in its own internal document, that the training data of its most profitable partner (OpenAI) is "theft," the situation is dramatic on its own — but the numbers are even harsher.
Microsoft's own data shows that since ChatGPT launched, click-through rates to some news publisher sites dropped by 83% to 93%; another batch of news organizations saw declines of 51% to 94%. In other words, the news industry is not just having its content stolen — its traffic is being stolen too. Hecht admits in the document that "almost no one wants their content used this way, and almost no one is paid for it" — but this sentence was meant for internal colleagues. The argument he really wanted to push is that this pipeline creates a "vicious cycle" that hurts both the model and the entire Web. When original content creators stop producing because they have no revenue and no traffic, the model's available corpus will also dry up.
Nadella admits in Congress vs Brockman says "cool" in Slack
The most dramatic part isn't the internal document — it's what came next. Microsoft CEO Satya Nadella testified before Congress and personally acknowledged that AI companies should not violate news sites' terms of service by bypassing paywalls. That's effectively repeating Brent Hecht's internal position in a public forum. But according to OpenAI internal information obtained by Ars Technica, when an OpenAI employee notified president Greg Brockman that "our crawler found a way to bypass the New York Times paywall," the reply was "cool."
One side's internal document calls it theft, another side's Congressional testimony admits it's a breach, and the side actually bypassing the paywall has its executives celebrating the bug in an internal Slack. This isn't a one-off mistake — it's a long-running tug-of-war among top companies that "admit the problem with words and keep scraping with actions." If even Microsoft, OpenAI's largest investor and compute supplier, admits "this shouldn't be done," how much defense space is left? (See the original Ars Technica report: Microsoft exec called AI scraping "the largest theft of labor in human history")
So the real signal from Hecht's document isn't the fact that "AI companies train on news" — it's that the cost of this has now been quantified, and quantified inside Microsoft's own internal PPT. A click-through decline range of 83% to 94% is not an external researcher's suspicion; it's the involved party's internal estimate. When such numbers first surface in public reporting through leaked internal documents, the discussion is no longer purely an ethical debate — it becomes factual groundwork that regulators and courts can use as evidence.