Daily Podcast full article
OpenAI copyright defense faces damaging filings
Newly unsealed court filings in the New York Times-led copyright case against OpenAI and Microsoft have sharpened the central fight over generative AI: whether large-scale training on protected journalism is transformative fair use, or an unlicensed substitution business built on publishers’ work.

A damaging new record for the fair-use fight
OpenAI and Microsoft entered the latest phase of the New York Times copyright litigation arguing that training artificial intelligence systems on news articles is lawful because the use is transformative and does not replace the original journalism. Newly public filings now put that defense under sharper pressure: according to reports on the unsealed material, the plaintiff publishers say internal statements from Microsoft and OpenAI executives show that the companies understood their products could substitute for news, reduce traffic to publishers and rely heavily on copied articles .
The case, pending in Manhattan federal court before U.S. District Judge Sidney Stein, began with the Times’ 2023 lawsuit and has since drawn in other news organizations, including the New York Daily News, The Intercept and the Center for Investigative Reporting . The current procedural moment is summary judgment, meaning the parties are asking the court to decide major issues before trial based on the evidentiary record . That makes the newly unsealed language unusually important: the dispute is no longer only about abstract copyright theory, but also about what company insiders allegedly said while the systems were being built and commercialized.
The publishers’ framing is blunt. They argue that OpenAI and Microsoft copied millions of articles without permission, used that material to train and improve commercial AI systems, and then deployed products that can answer users directly instead of sending them to the original news source . Microsoft and OpenAI continue to deny that this amounts to infringement, maintaining that AI training is protected by fair use and that their products do not unlawfully replace journalism .
The phrases plaintiffs want the judge to notice
The most politically explosive material concerns internal characterizations of AI training and market substitution. Reuters reported that court filings made public on September 17 say OpenAI and Microsoft executives privately described AI products as substitutes for journalism, a point the news organizations argue undercuts the companies’ fair-use defense . Bloomberg Law reported that OpenAI executive Nick Turley wrote that the company’s products were “largely substitutive,” while Microsoft CEO Satya Nadella testified that chatbot interactions had substituted for publisher sites, according to the publishers’ motion .
The filings also quote Microsoft director of applied science Brent Hecht warning that mass ingestion of others’ work could be perceived as an “astonishing theft” and possibly the largest theft of labor in history . Microsoft’s response is that Hecht was expressing one employee’s individual perspective, not the company’s legal position or an official admission . That distinction will matter: courts generally decide fair use as a legal question, not by treating loose internal language as dispositive. But the plaintiffs will likely argue that the comments illuminate market harm, knowledge and intent.
The Times report adds another thread: internal OpenAI concern that ChatGPT could draw users away from reading news elsewhere, including a June 2023 memo in which Turley described AI as an existential threat to publishers and a later note that AI products would become more substitutive as they improved . For publishers, those statements go directly to the fourth fair-use factor, which examines the effect of the use on the potential market for the copyrighted work. For OpenAI and Microsoft, the counterargument is that prediction, commentary or anxiety inside a company is not the same as proof of unlawful substitution.
Paywalls, datasets and the “AI pipeline”
Beyond the rhetoric, the unsealed material reportedly describes how publisher content entered the AI pipeline. TechCrunch reported that the filing discusses OpenAI and Microsoft allegedly obtaining content through mass scraping, using data derived from Common Crawl, building datasets such as WebText and WebText2, and relying on Microsoft-related initiatives described as Project Taxi and Project Mango . The same report says the plaintiffs allege Project Mango contained copies of at least 160,903 unique publisher works and that a Common Crawl-derived dataset included more than two million documents from the Times’ domain .
The paywall allegations may be especially sensitive. Reuters reported that OpenAI president Greg Brockman allegedly responded positively when told employees had found a way around the New York Times paywall . The Times likewise reported that the newly public documents quote an exchange in which a staffer described a paywall workaround and Brockman responded, “ah nice” . These claims, if accepted in context, could make the case feel less like ordinary web indexing and more like deliberate acquisition of protected content whose owner had tried to restrict access.
The publishers also say the copying was not incidental. AFP, carried by TechXplore, reported that the court document alleged OpenAI scraped content from more than 10 million articles, nearly a third from the Times alone . The plaintiffs are seeking summary judgment in their favor, while OpenAI and Microsoft argue that the training process transformed the material into new statistical and functional systems rather than republishing articles as competing copies .
Why fair use remains the hinge
The filings are damaging because they attack the cleanest version of the AI industry’s defense. In that version, training resembles reading: a model learns patterns from publicly accessible text, then generates new outputs rather than storing or reselling the original works. OpenAI and Microsoft have argued that their use is transformative and does not compete with or replace the journalism on which the systems were trained .
The publishers want to show the opposite. If an AI answer engine can satisfy a user’s need for information without a click to the article, and if company insiders anticipated that result, then the training and deployment may look less like research and more like building a substitute market. The Washington Post reported that the Times is asking the court to find liability for unauthorized copying at each stage of the AI pipeline . That phrase matters because it widens the question from final chatbot outputs to collection, dataset construction, model training, product integration and post-lawsuit mitigation.
The government’s posture adds complexity. Reuters reported that the Trump administration filed in support of OpenAI and Microsoft in early September, telling Judge Stein that AI training is extraordinarily transformative . AFP reported that the Justice Department brief invoked scientific progress, economic growth and national security . That support may reinforce the defense’s argument that the broader public interest favors allowing AI development, but it does not erase the publishers’ market-harm arguments or the specific allegations about paywalled and copied content.
Microsoft’s narrowing strategy
Microsoft’s public response is to narrow the meaning of the internal statements. The company says Nadella’s testimony is consistent with its legal position because he was speaking about broad changes in how people find and consume information, not conceding copyright liability . It also says Hecht’s comments were not legal analysis and do not represent Microsoft’s view .
That is a rational litigation posture. Companies often contain internal debate, worst-case forecasting and ethical disagreement. A court may discount colorful phrases if they are unsupported by the legal record. But the plaintiffs do not need every phrase to become an admission. They need the filings to help tell a coherent story: the companies knew journalism was valuable, knew AI answers could replace visits to news sites, and proceeded without the licenses publishers say were required.
OpenAI’s silence in several reports also leaves Microsoft carrying much of the quoted response in the public record . That does not necessarily signal weakness; defendants often limit comment while litigation is active. But in reputational terms, the absence of a detailed public rebuttal allows the publishers’ narrative to dominate this news cycle.
What is at stake for AI and news
The immediate question is whether Judge Stein will grant summary judgment to either side or leave key issues for trial. AFP reported that a ruling is not expected until 2027 . However the timing plays out, the newly public filings have already shifted the case from a technical dispute about training data to a broader test of the economic bargain between AI companies and professional publishers.
If the publishers prevail on core issues, AI developers could face higher data costs, stronger pressure to license archives and greater obligations to document exactly what was used in training. If OpenAI and Microsoft prevail, the decision could strengthen the industry’s argument that large-scale text ingestion is lawful when used to build general-purpose models. Either outcome will shape licensing markets, product design and the degree of provenance documentation expected from frontier AI companies.
For now, the filings do not decide the case. They do something more subtle: they make the fair-use defense harder to present as frictionless. The court must still apply copyright doctrine, but the public record now includes alleged internal warnings about substitution, paywalls and the sustainability of the news supply chain. In copyright court, that may be enough to turn a theoretical AI training case into a concrete fight over who pays for the raw material of information.
Sources from the last 72 hours
- [1]OpenAI, Microsoft executives’ quotes on AI training threaten copyright defense, news outlets argueSep 17, 2026, 9:47 PM UTC
- [2]Microsoft and OpenAI Workers Worry About ‘Largest Theft of Labor’ in HistorySep 17, 2026, 12:00 AM UTC
- [3]OpenAI, Microsoft Leaders Admitted AI Threatens News, Books (2)Sep 17, 2026, 6:03 PM UTC
- [4]Microsoft exec called AI the ‘largest theft of labor’ in history, court records showSep 17, 2026, 9:47 PM UTC
- [5]Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings revealSep 17, 2026, 7:46 PM UTC
- [6]NYT alleges Microsoft, OpenAI knew using news content was theftSep 18, 2026, 12:00 AM UTC
AI-generated article based on recent web research, then preserved as a dated editorial snapshot.

Comments
Be the first to comment.