Microsoft exec called AI scraping the “largest theft of labor in human history”

Microsoft and OpenAI are under fire not just for their latest AI rollouts but for the way they built them. The leaked language is now being wielded by the plaintiffs to argue that the firms knew they were crossing a line.

Sep 18, 2026 - 16:03
0 5
Microsoft exec called AI scraping the “largest theft of labor in human history”

Microsoft and OpenAI are under fire not just for their latest AI rollouts but for the way they built them. Leaked internal memos, revealed in a court filing by The New York Times and other news plaintiffs, show that senior engineers at both firms were aware that scraping news articles to train large language models was tantamount to “the largest theft of labor in human history.” The documents also lay out a stark prediction: the very products they were about to launch—ChatGPT, Copilot and other AI assistants—could cannibalize the news ecosystem that fed them. As the legal battle heats up, the leaked paperwork offers a rare glimpse into the calculus that guided the AI giants, and why the stakes matter to anyone who reads a headline on a phone.

Inside the memo that called news scraping “an astonishing theft”

Microsoft’s Director of Applied Science Brent Hecht repeatedly warned internal teams that the plan to scrape news content for AI training was “an astonishing theft of unprecedented proportions.” In one passage he went further, labeling it “perhaps the largest theft of labor in human history.” The phrasing appears in the same set of documents that the plaintiffs say contradict Microsoft’s public claim that using news articles falls under fair use. Hecht’s own note admits that “almost no one intended for content they created to be used in this fashion, nor are they compensated for its use.”

The memo also called the company’s fair‑use argument “a complete mockery of the idea of ‘fair use.’” By framing the practice as a moral and legal misstep, Hecht’s internal warning underscores a tension between the tech‑industry narrative of open data and the reality that newsrooms never signed up for their work to become fodder for AI. The leaked language is now being wielded by the plaintiffs to argue that the firms knew they were crossing a line.

OpenAI’s own alarm: “existential threat” to publishers

Across the corporate fence, OpenAI’s ChatGPT lead Nick Turley sent an internal message warning that publishers faced an “existential threat” from commercial AI products trained on their articles. Turley’s note framed the AI as a substitute for news providers, saying that users could get the same information without ever clicking through to the original site. The message also described a “doom loop” in which the AI’s reliance on news content would erode the very supply chain that fuels its usefulness.

One OpenAI engineer later echoed the sentiment, noting that “no matter how prominently we show the links, users won’t click.” Turley’s own assessment that chatbots are “largely substitutive, period” and will become “more and more substitutive as they get better” mirrors the very pattern the news industry fears: AI tools that answer questions directly, leaving the original article untouched and un‑monetized.

The numbers that back up the fear

Both companies’ internal data confirm that the feared traffic loss is not hypothetical. Microsoft’s own analytics show click‑through rate drops of 83‑93 percent for some plaintiffs and 51‑94 percent for others after AI products went live. Those figures line up with separate reports of low click‑through rates from ChatGPT search results, as well as news outlets’ own accounting of traffic declines. The convergence of these data points suggests a concrete impact on ad revenue and subscription renewals for the affected publishers.

For the newsrooms, the math is simple: fewer clicks mean less revenue, which in turn reduces the resources available to produce the high‑quality reporting that AI models rely on. The plaintiffs argue that this feedback loop—AI stealing clicks, which then starves the news ecosystem of the content that trains the AI—creates a “doom loop” that threatens both the media industry and the reliability of the AI itself.

Legal strategy: proving substitution, not just copying

The lawsuits hinge on more than verbatim copying; they focus on substitution. Plaintiffs contend that the AI outputs are “substantially similar” to the original articles and that the products are being used as direct replacements for news consumption. In practice, they tested the models by repeatedly asking for the “next line,” requesting bullet‑point summaries, and even prompting the bots to “rate the bias” of specific pieces. In many instances, the chatbots spat out long excerpts or near‑verbatim passages, especially when asked to summarize or rate bias.

Because the plaintiffs have limited their initial claim to articles that demonstrate “extensive verbatim overlap,” they aim to establish a clear pattern of substitution that would weigh against a fair‑use defense. The argument is that even if some excerpts fall under permissible use, the broader business model—training on massive swaths of news without compensation—creates a market harm that the courts cannot ignore.

Executive testimony and internal jokes about paywalls

During the trial, Microsoft CEO Satya Nadella testified that AI firms should not be sidestepping paywalls or violating site terms of use. He described the AI experience as “giving you the information right there on the website on the AI platform versus needing to go to the underlying source,” effectively acknowledging that the chatbots siphon clicks away from news sites.

OpenAI’s internal communications reveal a more casual attitude toward the same issue. When a staffer named Nick Ryder flagged a “hack” that let OpenAI crawlers bypass the New York Times paywall, President Greg Brockman responded with a terse “Ah, nice.” The exchange, captured in the court documents, suggests that the technical work to evade paywalls was not only known but also approved at the highest levels.

Company defenses: fair use versus transformation

Microsoft’s public response frames its AI products as “transformative fair use” that do not substitute for news sites. A spokesperson emphasized that Nadella’s remarks were “broad principles and changes underway in how people find and consume information,” and should not be read as a legal conclusion. The spokesperson also downplayed Hecht’s memo, calling it “one employee’s individual perspective” and not reflective of company policy.

OpenAI, for its part, has not responded to the request for comment. The silence leaves the leaked internal messages as the primary evidence of the firms’ awareness of the substitution risk and the steps taken to mitigate—or ignore—paywall protections.

What this means for everyday readers

If the courts side with the news publishers, the immediate effect could be tighter restrictions on how AI models are trained, potentially slowing the rollout of new features in ChatGPT, Copilot and other Microsoft‑backed tools. For users, that might translate into fewer “instant answers” and a return to more traditional news‑site visits for in‑depth coverage.

Beyond the courtroom, the dispute spotlights a broader question about the digital commons: who gets to profit from the data that fuels AI? The plaintiffs argue that preserving incentives for journalists is essential to a “healthy society,” a point that resonates far beyond the headline‑click metrics. As AI continues to reshape information consumption, the outcome of this case could set a precedent for how tech giants negotiate the balance between innovation and the labor that underpins it.

This article was produced with AI-assisted research and editorial support. Reporting is based on the source material cited below. Sources: Ars Technica; arstechnica.com; Global1.News (18 September 2026).

By Nova Chen, Staff Writer

What's Your Reaction?

Like Like 0
Dislike Dislike 0
Love Love 0
Funny Funny 0
Wow Wow 0
Sad Sad 0
Angry Angry 0
Nova Chen

Trend Reporter at Global1.News. Based in San Francisco, tracking the stories crossing from social platforms, forums, and community discussions into mainstream news — tech breakthroughs, cultural shifts, and world events that real people are engaging with right now.

Comments (0)

User