How do AI systems such as ChatGPT, Claude and Gemini appear to know so much about the world? They are trained on enormous volumes of information collected from the internet. However, some of that material may not have been freely available for such use. OpenAI, Anthropic and other AI companies have faced allegations from authors, publishers and news organisations that they used copyrighted material to train their models without permission or compensation. There have also been concerns within the companies themselves. One Microsoft executive described AI scraping as the “largest theft of labour in human history.”
The remark came from Brent Hecht, Microsoft’s director of applied science. His comments became public through unsealed court documents in the copyright lawsuit filed by The New York Times and other news organisations against Microsoft and OpenAI.
In the filing, Hecht discussed the use of internet content for training AI models and characterised the practice as the biggest theft in history.
“Millions of people around the world will soon consider large models ‘hoovering up’ all their work to be an astonishing theft of unprecedented proportions,” he wrote. He also pointed out that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”
Microsoft and OpenAI, however, have argued that using copyrighted content to train AI models falls under fair-use protections. The companies maintain that their systems transform the material instead of simply reproducing it. The New York Times and other publishers dispute this, arguing that AI products can effectively replace the original news sources.
The court filings also shed light on how AI chatbots are changing the way people consume news. Rather than visiting a news website and clicking through articles, users can ask an AI chatbot for information and receive an answer directly on the platform. During his deposition, Microsoft CEO Satya Nadella acknowledged this shift, saying that interacting with chatbots has “substituted giving you the information right there on the website on the AI platform versus needing to go to the underlying source.”
News organisations argue that this creates a challenging cycle for publishers. AI models depend on material produced by news outlets for their training, while AI chatbots can simultaneously reduce the number of people who visit those same websites. The filing refers to Microsoft data showing that click-through rates to The New York Times and Daily News domains were 83 to 93 per cent lower through Copilot’s “answer engine” than through conventional Bing Search.
The documents also quote OpenAI’s ChatGPT chief Nick Turley describing the products as “largely substitutive” and saying they would become “more and more substitutive as they get better.”
The filings contain further internal assessments of AI models’ ability to process news content. OpenAI co-founder Greg Brockman wrote that the models were “particularly good at predicting text of news articles” and “very good at any news task”. He made the observations while discussing the models’ performance on content from The New York Times.
The New York Times filed its lawsuit against Microsoft and OpenAI in 2023, roughly a year after ChatGPT was launched. Other news organisations, including the New York Daily News, The Intercept and the Center for Investigative Reporting, subsequently became part of the case.
Microsoft and OpenAI continue to rely on their fair-use argument, while the news organisations involved in the lawsuit are seeking a court ruling in their favour on their copyright claims.
