When OpenAI and Microsoft were building the systems that would become some of the most widely used artificial intelligence products in history, their own internal documentation warned them they were doing something deeply corrosive. The Verge has reported on recently unsealed court documents from the New York Times' ongoing lawsuit against the two companies, revealing that employees characterized their data-scraping practices as potentially triggering a "doom loop" for the web and, in a phrase that will be difficult to walk back, described the operation as the "largest theft of labor" in some relevant framing.
The significance of those words cannot be overstated. In litigation of this kind, what a company's own people wrote in internal documents carries a weight that external criticism never can. Plaintiffs spend years trying to prove that defendants knew their conduct was harmful. Here, according to The Verge's reporting, the defendants appear to have written it down themselves.
To understand why this matters, it helps to remember how the large language model industry justified its early appetite for data. The standard argument, deployed repeatedly in response to lawsuits from authors, visual artists, musicians, and now major publishers, was essentially one of transformation: that training a model on existing work is categorically different from reproducing or displacing that work, and that the output is something genuinely new. That argument is still being litigated in multiple jurisdictions, and it may yet succeed in some form. But it depends, at least partly, on the premise that the companies engaged in this practice without a clear-eyed understanding that they were taking something of value from someone else without compensation. Internal documentation describing the activity in terms of theft and systemic damage to the web's underlying economy makes that premise considerably harder to maintain.
The "doom loop" framing refers to a real and well-understood dynamic. Much of the web's content is produced because search engines surface it to readers, which generates traffic, which generates revenue, which funds more content. If AI systems begin answering questions directly rather than routing users to source material, that traffic collapses. Publishers lose the economic basis for employing journalists, researchers, and writers. The pool of new, reliable, human-generated content shrinks. The models, which depend on fresh data to remain accurate and current, then have less quality material to learn from. The loop feeds on itself. That this dynamic was identified inside OpenAI and Microsoft suggests it was not an unforeseen side effect but a foreseeable consequence that the companies chose to accept.
The New York Times filed its lawsuit at the end of 2023, and it has always been among the more legally credible of the wave of AI copyright cases, partly because the Times has the resources to litigate seriously and partly because it could point to specific, verbatim reproductions of its journalism in model outputs. But the unsealed documents described by The Verge shift the terrain somewhat. The legal question of whether training on copyrighted material constitutes infringement is still genuinely unsettled. The reputational and regulatory question of whether these companies acted in good faith is now considerably more complicated.
For Microsoft, the exposure is particular in one respect. The company made an enormous financial commitment to OpenAI at a moment when these concerns were apparently already being documented internally. That raises questions for investors, for regulators examining the partnership's competitive implications, and for the broader technology press that largely celebrated the deal as visionary. For OpenAI, which has spent recent months pursuing a restructuring into a for-profit entity and courting additional investment at a valuation in the hundreds of billions of dollars, the reputational damage from internal documents using the language of theft is the kind of thing that complicates a public narrative built around the idea of building AI safely and responsibly.
The consequences extend beyond the two defendants. Every major AI company that trained models on scraped web content has some version of this exposure. If courts or regulators decide that internal awareness of harm is legally or commercially relevant, the industry's collective liability calculus changes. Publishers and creators who have been settling, waiting, or simply watching may feel newly emboldened.
The immediate thing to watch is how OpenAI and Microsoft respond to the substance of these documents becoming public, and whether their legal teams attempt to recontextualize the language or dispute The Verge's characterization of what the documents show. Beyond that, the trajectory of the Times lawsuit itself remains the most consequential single case in AI copyright law, and these documents are likely to feature prominently as it moves toward whatever resolution awaits. More broadly, any AI company that has not yet audited its own internal communications for similar candor might want to do so before the next set of court filings arrives.




