Tech • AI • Robotics • Game

VIDEO
ENFR

Daily Podcast full article

OpenAI faces the copyright dossier

Newly unsealed filings in the New York Times-led copyright case have turned OpenAI and Microsoft’s fair-use defense into a fight over their own words: plaintiffs say executives described AI products as substitutes for journalism, anticipated harm to publishers, and still built the training-data economy around uncompensated news content.

Generated September 18, 2026 at 10:35 AM UTC1289 words

A copyright case becomes a business-model case

The copyright fight around OpenAI is no longer only about whether a model can learn from protected articles. It is now about whether OpenAI and Microsoft knew that the products built from those articles could replace the publishers that created them. Court filings made public on September 17 say executives and employees at both companies privately described AI systems as substitutive, powerful at news tasks and dangerous to the economics of journalism, according to fresh reporting on the unsealed record .

That matters because the companies’ central defense has been fair use. OpenAI and Microsoft argue that training on millions of newspaper articles transforms protected text into a new technological product and does not compete with the original journalism . The news plaintiffs, led by the New York Times and joined by other outlets, say the newly visible portions of the record attack that defense at its most vulnerable point: market substitution .

The words plaintiffs wanted the court to see

The filings point to statements attributed to senior figures including Microsoft CEO Satya Nadella, OpenAI president Greg Brockman, OpenAI’s head of ChatGPT Nick Turley, and Microsoft director of applied science Brent Hecht . Reuters reported that the plaintiffs quoted Turley as saying publishers face an “existential threat” from products that are “largely substitutive” and expected to become more substitutive as they improve . Bloomberg Law likewise reported that Nadella testified that chatbot conversations had “substituted” for publisher sites, while Turley wrote that OpenAI products were largely substitutive .

The most explosive line comes from Hecht. According to Reuters, the filing quoted him as saying that many people would view AI companies “hoovering up” their work as an “astonishing theft of unprecedented proportions” . Bloomberg Law reported an even sharper version from the unsealed material: Hecht said copying millions of news articles without permission could be the “largest theft of labor in human history,” and that a fair-use win for the companies could make a “complete mockery” of fair use .

For the publishers, these are not stray comments. They are being used as evidence that the defendants internally understood the same market harm they deny in court. For Microsoft, the comments do not decide the legal issue. A company spokesperson said Nadella’s testimony addressed broad changes in how people find information and should not be confused with legal conclusions about copyright, and said Hecht’s comments reflected one employee’s view rather than Microsoft’s position .

Paywalls, scraping and the “excellent at news” problem

The filings also revive a practical question: if a news article is behind a paywall, can it become training material without a license? Reuters reported that the outlets cited Brockman as writing that large language models were especially good at predicting news article text and “excellent at news,” and that he responded approvingly when employees discovered a way around the New York Times paywall . TechCrunch reported that the unsealed material alleges paywall bypassing, mass scraping and removal of copyright notices from training data, while noting that much of the new information comes from the Times’ brief and that underlying exhibits remain sealed .

That caveat is important. The public is seeing plaintiffs’ characterizations and selected quotations, not the full universe of discovery. But in litigation, especially at the summary-judgment stage, carefully selected internal statements can shape the judge’s understanding of what disputes must go to trial. The plaintiffs want Judge Sidney Stein to see not a neutral act of machine learning, but a chain: copying articles, building products that answer news questions, reducing the need to click through, and weakening the publishers that supplied the raw material.

The “doom loop” theory

The strongest economic argument now emerging is the “doom loop.” Ars Technica reported that Microsoft and OpenAI documents, as described by the plaintiffs, reflected a fear that AI products could damage the news organizations on which they depend for reliable information . According to Ars, one Microsoft document described a loop that would hurt model performance and the web at the same time, because the end product threatens the economic foundations of its own content suppliers .

The traffic figures reported from the filings are central to that claim. TechCrunch reported that Microsoft data showed Copilot’s answer engine caused click-through rates for the New York Times domain to fall by as much as 93% compared with traditional Bing search . Bloomberg Law reported claimed click-through drops of 83% to 93% for Times and Daily News content and 51% to 94% for Ziff Davis domains . If those figures are accepted, the plaintiffs can argue that the harm is not speculative: AI answers do not merely learn from journalism; they intercept demand that previously sent readers to publishers.

That is why the case threatens more than one lawsuit. The training-data economy depends on the idea that large-scale ingestion is lawful, or at least negotiable after the fact. If the court finds that news training plus substitutive outputs weighs against fair use, the price of high-quality data could rise sharply. If the court accepts OpenAI and Microsoft’s transformative-use argument, publishers will face pressure to license on weaker terms or build technical walls that may be hard to enforce.

A broader front of publishers

The dispute is also widening. Bloomberg Law reported on September 16 that around 160 regional, local and trade news publications owned by 26 publishers filed a new complaint in the Southern District of New York against OpenAI and Microsoft . The group reportedly includes outlets such as the Tampa Bay Times, the Austin Chronicle and Florida Trend, and accuses the companies of building valuable AI businesses on publishers’ work without permission or compensation .

That broadening is strategically significant. The New York Times is large enough to litigate alone, but local publishers bring a different story of harm. They operate with thinner margins, depend heavily on referral traffic and often produce civic information that is expensive to gather but easy for an AI interface to summarize. If their claims consolidate around substitution, licensing and loss of clicks, the court will be looking not only at elite national journalism but at the fragile economics of the local information supply.

Fair use still needs a compiler

OpenAI and Microsoft are not without legal support. Reuters reported that the Trump administration told the court on September 1 that AI training is “extraordinarily” transformative . The companies argue that training does not simply republish articles; it extracts statistical patterns to build systems capable of generating new outputs . That is the core of the fair-use case for the AI industry.

But fair use is a balancing test, not a magic keyword. Courts look at purpose, nature, amount and market effect. The newly unsealed comments are aimed above all at purpose and market effect: were the defendants building a new tool that complements journalism, or a substitute that consumes journalism and then competes with it? The plaintiffs want the court to treat internal warnings as proof that the companies anticipated the damage. The defendants want those comments treated as observations, individual opinions or product-market discussion rather than admissions of infringement.

The next phase is therefore less about slogans than architecture. The court must decide whether the legal system can compile a rule for AI training that distinguishes learning from substitution, indexing from replacement and transformation from extraction. Until then, OpenAI’s copyright dossier remains the test case for a larger question: whether the web’s most valuable text can be treated as free infrastructure for models that may redirect the web’s own revenue.

Comments

Be the first to comment.

Sources from the last 72 hours

  1. [1]OpenAI, Microsoft executives’ quotes on AI training threaten copyright defense, news outlets argueSep 17, 2026, 9:47 PM UTC
  2. [2]OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2)Sep 17, 2026, 6:03 PM UTC
  3. [3]Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings revealSep 17, 2026, 7:46 PM UTC
  4. [4]Microsoft exec called AI scraping the “largest theft of labor in human history”Sep 17, 2026, 8:10 PM UTC
  5. [5]Local Outlets Join Copyright Fight Against OpenAI, Microsoft (1)Sep 16, 2026, 8:34 PM UTC

AI-generated article based on recent web research, then preserved as a dated editorial snapshot.