Microsoft memos warned OpenAI training risked "largest theft of labor"
In brief
- Microsoft employees warned OpenAI's scraping of paywalled articles amounted to largest labor theft in human history
- Internal 2023 memos stated AI models 'hoovering up' people's work would 'destroy their supply chain'
- CEO Satya Nadella testified paywalled content requires licensing; he'd have forced model retraining if aware
- OpenAI staffers discussed bypassing New York Times paywall; executives defended training as fair use
Microsoft's Internal Warnings
An internal Microsoft document from 2023 stated that millions of people around the world would soon consider large models "hoovering up" all their work to be theft of unprecedented proportions. The same memo warned that large AI models "are a product that destroys its supply chain." Microsoft attributed these memos to Brent Hecht, a director of applied science who also held a Northwestern University post, and said they do not represent company views. The company added that Hecht was not a decision maker and was employed to "present divergent and asymmetric perspectives."
Nadella's Testimony and OpenAI's Practices
Satya Nadella, Microsoft's chief executive, testified that "anything that is paywalled should be licensed by anyone who wants to use it." He also said that had he known OpenAI was training on paywalled content, he would have exercised Microsoft's right to make it retrain its models. A spokesman clarified that Nadella "spoke to broad principles" regarding how people find and consume information.
At OpenAI, the internal friction ran deeper. A staffer told president Greg Brockman about building a "hack" to bypass the New York Times paywall. Brockman replied: "ah nice."
Nick Turley, who ran the ChatGPT team, wrote in June 2023 that AI posed an "existential threat" to publishers, and later noted in February 2024 that AI products "will get more and more substitutive as they get better." An OpenAI engineer wrote in February 2023 that "no matter how prominently we show the links, users won't click."
The Broader Critique
Jack Clark, then-policy director at OpenAI, warned in a 2020 memo to Brockman and Sam Altman that the company was "creating systems that substitute for the labor of the people that define the 'culture' of society," and would "become the symbol of how Silicon Valley is thoughtlessly stepping into other parts of life and leaving a mess on the carpet." Clark later left to co-found Anthropic.
Both Microsoft and OpenAI argue the training was fair use, transforming articles into new work rather than substituting for the originals. Steven Lieberman, representing the New York Daily News and seven other papers, countered that "the world can see what OpenAI and Microsoft thought all along about the fairness of their own behavior."
The New York Times brought suit against both companies in late 2023, since joined by eleven other publishers. The case has already forced OpenAI to preserve 20 million ChatGPT conversation logs. Judge Sidney Stein of the Southern District of New York is weighing summary judgment motions. OpenAI has contested the claims throughout.
Frequently asked questions
What did Microsoft employees warn about regarding AI training?
Internal Microsoft memos from 2023 warned that large AI models 'hoovering up' people's work constituted theft of unprecedented proportions and could trigger a doom loop degrading model quality. Employees questioned whether OpenAI's scraping of paywalled news amounted to the largest theft of labor in human history.
What is OpenAI's defense against the copyright lawsuit?
Both Microsoft and OpenAI argue the training constituted fair use, transforming articles into new work rather than substituting for originals. OpenAI has contested the claims throughout the lawsuit brought by the New York Times and eleven other publishers.
What did OpenAI executives say internally about the training?
Nick Turley, who ran ChatGPT, called AI an existential threat to publishers and noted AI products become increasingly substitutive. An engineer acknowledged users won't click paywalled links. A staffer even discussed building a hack to bypass the New York Times paywall.


