Click to open contact form.
Your Global Partners in the Business of Innovation

U.S. Department of Justice Says LLM Training on Copyrighted Works Is Fair Use

Client Updates / September 29, 2026

Written by: Haim Ravia, Dotan Hammer

On September 1, 2026, the United States filed a Statement of Interest in the copyright infringement litigation that the New York Times asserted against OpenAI in the Southern District of New York. Although directed at the New York Times dispute, it is framed to apply to all parties, including book authors and publishers.

The Government’s central submission is that training a large language model on copyrighted written works is “exceedingly transformative” because the copying is undertaken “not to duplicate the work’s expressive content, but as part of a process to learn from and act on statistical patterns in written text.” That purpose differs fundamentally from the entertainment or education of an audience for which the original works were created. The adverse impact of the use on commerciality, which is the fourth fair use factor, is treated as insignificant on the principle that the more transformative the use, the less weight attaches to other factors. The second and third factors, the Government submits, are “unlikely to play a major role,” distinguishing the making of copies during training from making content “accessible to a public for which it may serve as a competing substitute.”

The argument that matters most in practice is the insistence on a use-by-use analysis that separates training from outputs. Training, the Government contends, “simply creates a copy of a protected work in order to teach an LLM to recognize relationships between data,” and does not reveal a significant amount of original authorial expression to the public, thus producing no cognizable substitution. Whatever legal questions particular outputs may raise, “that should not bear on the transformative nature (or any other aspect) of the training use.” On that footing, the U.S. Government attacks the opposing theory that merges training and outputs into a single use, saying that it overlooks “arguably the most basic element of copyright infringement—whether the works are substantially similar.” Cognizable harm, the Government maintains, requires “significant substitutive competition, not competition generally.” According to the Government, genre-level competition is not infringement. The filing also invokes the public-benefit dimension of the fourth factor, observing that New York Times writers themselves use LLMs to conceptualize and edit articles and that independent publishers can use such tools to level the playing field.

National security supplies the asserted federal interest. The filing cites a Government Accountability Office warning that failure to adopt and effectively integrate AI could hinder national security. Rules restricting AI development in the United States are said to “give a competitive advantage to foreign adversaries who are not so encumbered”. The Government opposes a licensing requirement as an anticompetitive barrier that would favor large legacy media companies over startups, while expressly declining to take a position on whether a licensing regime would be financially or logistically feasible.

Click here to read the Statement of Interest of the United States on fair use.

MEDIA HIGHLIGHTS