Tech giants admit to ‘largest theft of labour in human history’
Image: Valent Lau @valentlau via Unsplash
When Adolph Ochs bought the New York Times out of bankruptcy in 1896, he promised readers impartial news “without fear or favour, regardless of any party, sect, or interest involved”.
In 2023 the newspaper sued Microsoft and OpenAI for copyright infringement. In court filings, it pointed out Ochs’s words “still animate The New York Times today, nearly two centuries later”, but warned its mission was once again in grave peril because AI companies were ripping off its proprietary content.
The Times leads a group of news publishers that are suing the tech giants for billions of dollars in damages for using their content without permission to train large language models and generate verbatim copies and detailed summaries mimicking the same “expressive style”.
The Times says when it tried to negotiate a deal with Microsoft and OpenAI for commercial use of its intellectual property, it was confronted with the “fair use” argument, which holds that generative AI models transformed copyrighted content into new products. “But there is nothing ‘transformative’ about using The Times’s content without payment to create products that substitute for The Times and steal audiences away from it,” the newspaper pointed out.
Since then, the case has yielded a treasure trove of internal company documents and anonymised ChatGPT logs – most recently through the release of newly unredacted filings that were first shared on X by Jason Kint, CEO of the trade association Digital Content Next.
The filings contain a series of astonishing admissions by tech executives about the impact their wholesale pillage is having on original content creation.
Microsoft’s director of applied science, Brent Hecht, described the predatory practice as “an astonishing theft of unprecedented proportions” and perhaps the “largest theft of labour in human history”. Hecht pointed out that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use”.
In another internal document, Microsoft warned its “AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time” by stealing content, then diverting users from visiting the sites they stole it from. Microsoft recorded a drop of 83-93% in click-through rates to New York Times articles. “It is highly unusual that an end-product threatens the economic foundations of its essential suppliers,” Microsoft concluded. “But that is the situation we have created for our LLM business with respect to its ‘content supply chain’.”
The tech executives have no illusions about the cost to writers, journalists and news publishers. Hecht points out that AI systems do not “distribute economic value down the supply chain”, admitting this “necessarily threatens the economic stability of those who create the content”. In the same vein, OpenAI Policy Director Jack Clark concedes the company’s “work on AI and creativity is going to increasingly lead to us creating systems that substitute for the labour of the people that define the ‘culture’ of society”. Microsoft, in another internal document, admits to a “real risk” that generative AI could “significantly disrupt the employment of the very people who generated the data on which the foundation model was trained”.
OpenAI’s head of ChatGPT, Nick Turley, goes even further, admitting news publishers faced an “existential threat” from generative AI. Despite knowing this, the company pressed ahead, “deeply motivated by the gazillions” it could make, as co-founder Greg Brockman put it. The court filing that the AI companies wanted to keep sealed also show how they tried to sneak around paywalls when scraping websites. When Brockman was informed about “a hack to get around nytimes paywall”, he replied: “Ah nice.”
David Buttle of the Standards for Publisher Usage Rights Coalition in the UK says these revelations should serve as a wake-up call to the publishing industry to take concrete steps to protect themselves from having their content looted. He has called for transparent, industry-wide machine-readable standards that are subject to external audit “to track and disclose exactly how AI companies are using the intellectual property”.
The illusion that AI companies can be trusted to act in good faith has been decisively shattered.