Microsoft Staff Asked Whether AI Scraping Was the 'Largest Theft of Labor in Human History'
Key Takeaways
- •Unsealed filings from the Times' 2023 copyright lawsuit show Microsoft employees privately compared OpenAI's use of news articles to an unprecedented theft of human labor and warned that training on AI-generated material could progressively degrade model quality.
- •Microsoft CEO Satya Nadella testified that paywalled content should be licensed and said he would have exercised Microsoft's right to force OpenAI to retrain its models had he known it was using paywalled material.
- •Internal OpenAI communications include a staffer's description of building a workaround for the Times' paywall, which drew a positive reply from president Greg Brockman.
- •Internal OpenAI findings, including an engineer's observation that users rarely click cited links, undercut the companies' argument that chatbots drive traffic back to publishers.
- •Both defendants maintain that training on news articles constitutes fair use as Judge Sidney Stein considers summary judgment motions in a case the Times filed in late 2023 and eleven other publishers have since joined.

Microsoft employees debated whether OpenAI's use of news articles amounted to "the largest theft of labor in human history" and could set off a "doom loop" that degraded the very models they were building, according to court documents unsealed Thursday and reported by the New York Times.
The filings stem from the copyright lawsuit the Times brought against OpenAI and Microsoft in late 2023, a case since joined by eleven other publishers. OpenAI has contested the claims throughout, and the litigation has already compelled the company to preserve 20 million ChatGPT conversation logs. Judge Sidney Stein of the Southern District of New York is weighing summary judgment motions, and documents are being unsealed as he considers them. Summary judgment allows a judge to resolve a case without a trial when no genuine dispute of material fact exists, which is why the unsealed record — including internal communications from both companies — carries weight at this stage of the litigation.
Internal Microsoft documents from 2023 anticipated the backlash. "Millions of people around the world will soon consider large models 'hoovering up' all their work to be an astonishing theft of unprecedented proportions," one memo said. The same author wrote that large AI models "are a product that destroys its supply chain." The remark pointed to a feedback loop in which models trained heavily on AI-generated material would progressively degrade their own output — the "doom loop" employees referenced.
Microsoft said the memos were written by Brent Hecht, a director of applied science who also held a post at Northwestern University, and do not represent company views. In a filing, the company said Hecht was not a decision maker and had been employed to "present divergent and asymmetric perspectives."
"Leaving a mess on the carpet"
Microsoft CEO Satya Nadella testified that "anything that is paywalled should be licensed by anyone who wants to use it," adding that had he known OpenAI was training on paywalled content, he would have exercised Microsoft's right to force it to retrain its models. A company spokesman said Nadella "spoke to broad principles" about how people find and consume information.
At OpenAI, a staffer told president Greg Brockman about building a "hack" to bypass the Times' paywall. Brockman replied: "ah nice."
Nick Turley, who ran the ChatGPT team, wrote in June 2023 that AI posed an "existential threat" to publishers. In February 2024, he wrote that AI products "will get more and more substitutive as they get better," noting elsewhere that AI "products are largely substitutive, period."
An OpenAI engineer observed in February 2023 that "no matter how prominently we show the links, users won't click" — a finding that cuts against the argument that chatbots send traffic back to publishers, a long-standing source of readers and revenue for news organizations.
In a 2020 memo to Brockman and CEO Sam Altman, then-policy director Jack Clark warned the company was "creating systems that substitute for the labor of the people that define the 'culture' of society," and "become the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet." Clark later left to co-found Anthropic, which referred a request for comment from the Times to OpenAI.
Both Microsoft and OpenAI argue that training on news articles constitutes fair use, transforming the material into new work rather than substituting for the originals. Fair use permits limited use of copyrighted material without a license, and whether large-scale training on published journalism qualifies is the central legal question the case will test.
"The world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior," said Steven Lieberman, an attorney representing the New York Daily News and seven other newspapers in the case.
The Times, itself a plaintiff, declined to comment to its own reporters, who said OpenAI did not respond to requests for comment. With summary judgment motions still pending, the next markers in the case are Judge Stein's ruling and whatever else the unsealing process brings into the open.