Microsoft's Own Exec Called AI News Scraping the Biggest Theft in Human History. Turns Out They Knew.
For years, Microsoft and OpenAI have fought tooth and nail to keep certain internal documents away from public scrutiny. The AI copyright litigation brought by major news organisations, led by The New York Times, has been grinding through the courts, with both companies insisting their training practices constitute fair use. They were also insisting, quietly but firmly, that nobody should see the paperwork.
That strategy has now collapsed. A motion for summary judgment was unsealed this week, and the documents inside it are awkward reading for both firms.
The standout: Microsoft's own Director of Applied Science, Brent Hecht, described scraping news content for AI training as 'an astonishing theft of unprecedented proportions' and floated the possibility it represented 'the largest theft of labor in human history.' In a separate document, Hecht suggested the practice made 'a complete mockery of the idea of fair use.' These are not the words of a hostile external critic. This is a senior technical employee at Microsoft.
Hecht also acknowledged, plainly, that 'almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.' Hard to square that with the fair use defence Microsoft is currently running in court.
Over at OpenAI, ChatGPT product head Nick Turley warned internally that publishers faced an 'existential threat' from commercial AI products capable of substituting for news providers entirely. One Microsoft document coined the term 'doom loop' to describe how AI products would hollow out the news ecosystem that generated the training data keeping those same products useful. 'It is highly unusual that an end-product threatens the economic foundations of its essential suppliers,' the document read, 'but that is the situation we have created for our LLM business with respect to its content supply chain.'
This was not abstract theorising. Microsoft's own data shows click-through rates for some news plaintiffs dropped between 83 and 93 percent. Others saw drops between 51 and 94 percent. Combined with independently reported declines in referral traffic from ChatGPT search, the doom loop appears to be functioning exactly as predicted.
There is also the small matter of OpenAI actively circumventing paywalls. When a staffer named Nick Ryder told OpenAI President Greg Brockman that engineers had found 'a hack' allowing their crawlers to bypass the NYT paywall, Brockman's response was: 'Ah, nice.' Meanwhile, Microsoft CEO Satya Nadella testified under oath that AI companies should not be violating news sites' terms of service. Make of that contrast what you will.
Nadella himself acknowledged during testimony that chatbots substitute for news platforms by delivering information directly rather than directing users to original sources. OpenAI's Turley agreed there was 'no good reason to click' when a chatbot has already handed over the answer. A software engineer at OpenAI said the same thing more bluntly in an internal message: 'no matter how prominently we show the links, users won't click.' Turley went further, describing chatbots as 'largely substitutive, period,' and predicting they would get more substitutive as the technology improved.
News organisations tested this in court by going beyond the earlier tactic of prompting chatbots to reproduce articles line by line. They found that requests for summaries, bullet points, bias ratings, or even 'pick an article from this site's homepage' could all produce lengthy verbatim excerpts. Their motion only seeks judgment on articles where outputs show extensive word-for-word overlap, because they believe the substitution evidence alone is sufficient to demolish the fair use argument.
Microsoft's response to all of this is that Hecht's comments 'reflect one employee's individual perspective' and 'do not represent the company's views.' Nadella's testimony, they claim, was just broad observation about changing media habits rather than any concession relevant to copyright law.
Steven Lieberman, counsel for the New York Daily News and seven sister papers, is not buying it. 'The evidence revealed here for the first time shows that OpenAI and Microsoft knew that what they were doing was wrong,' he told us. 'Throughout this case, defendants insisted these documents be treated as confidential so the public could not see them. Well, now the cat is out of the bag.'
The internal documents also allege that rather than address Hecht's concerns, Microsoft created a filter that he described as potentially an 'accidental cover up,' because it would reduce news groups' ability to see what content had been used for training. The firms are also accused of violating industry norms by selling a dataset originally purchased for Bing to OpenAI for model training, without informing or consulting the news organisations whose content it contained. OpenAI separately obtained a dataset of 1.8 million NYT articles from a third party that was contractually prohibited from using it for commercial purposes. OpenAI employees apparently recognised this was inappropriate and used the data anyway.
The news groups' argument is not just about compensation. It is about structural incentives. If courts rule that AI firms can freely train on published content without licensing it, there is no mechanism to break the doom loop. Every individual AI company is better off taking content for free while hoping others pay. That is a classic prisoner's dilemma, and right now nobody is blinking.
As evidence that the money is the point, the motion highlights Brockman writing that he was 'deeply motivated by the gazillions' to be made from commercialising OpenAI's technology. Subtle.
The news plaintiffs say they are ready for trial. Given what has just been unsealed, that confidence is at least understandable.