The U.S. government stepped into the debate over the application of fair use to AI companies using copyrighted works for LLM training. In a Statement of Interest filed September 1, 2026, in the consolidated OpenAI copyright litigation, the DOJ argued in favor of a broad application of fair use and that a contrary ruling would hand a competitive advantage to foreign rivals and choke off scientific progress.

A Statement of Interest, under 28 U.S.C. § 517, allows the government to weigh in on any case touching the government’s interest without leave of court. The government’s statement here addressed “the question whether the use of copyrighted works at the training stage—by copying works in order to feed data into the model as learning material—constitutes fair use.” In short, the government’s position is that the “training of AI models on copyrighted material, in and of itself, does not violate copyright laws.”

Contextualizing the fair use argument around policy, the government argued that copyright law’s purpose is not to reward authors’ labor but to promote scientific and artistic progress. The government notes a policy tension in copyright law: “[g]iving authors exclusive rights encourages the creation and dissemination of expressive works; limiting those rights ensures that secondary users are permitted fair breathing room and latitude to facilitate further expression.” It also emphasized the national-security importance of developing a robust domestic AI industry. It added that a narrow application of fair use would benefit larger LLM companies and established media companies. In particular, the government leans on the irony that the New York Times itself uses LLMs.

Turning to the fair-use factors, the government revisited the conflicting Bartz and Kadrey decisions, favoring the former over the latter.

Factor 1 (transformative use). The government emphasized the importance of the first fair-use factor (purpose and character) rather than the fourth factor (market effect). It raised a similar transformative-use argument that the Bartz court articulated. Specifically, the government argued, “The purpose of the copying (to build an intelligent, interactive model) differs in kind from the purpose of the copied work (to use language to directly entertain or educate a reading audience).” The government analogized the current LLM situation with the Supreme Court’s decision in Google LLC v. Oracle Am., Inc., 593 U.S. 1, 18 (2021). Google had copied copyrighted code but it had done so to operate within a new smartphone environment that rendered the purpose distinct from the original. The Court held that copying facilitated new expression and that a ban on such copyright would impede Google’s “ability to exploit its own creative expression.” The government posited that the current AI-training-data issue presents an even stronger case for fair use.

Factor 4 (market effect). On the fourth fair-use factor, the government argued that use for LLM training does not act as a substitute for copyrighted work. The government distinguished “whether the use could cause ‘some loss of sales’” from whether there is “significant substitutive competition.” The government argued that “[u]sing a copyrighted work to train an LLM—without more—generally does not result in this sort of substitution because it does not “reveal[]” a significant amount of original “authorial expression.” The government disagreed with the Kadrey court’s fourth-factor analysis and theory of “market dilution.” It argued that Kadrey improperly collapses LLM training and outputs as a single use. The government also faulted Kadrey for focusing on the economic harm from LLMs rather than the public benefit from the development of LLMs.

Factors 2 and 3. Buried in a footnote, the government also argues that the nature of the work (Factor 2) and the amount used (Factor 3) both support fair use because training does not make copied works “accessible to the public” as competing substitutes.

Outputs: A central theme of the Statement is that courts should analyze AI training and AI outputs as distinct uses for copyright purposes. Consistent with that view, the government acknowledged that its Factor 1 arguments address the training stage, not the output stage. It nevertheless argued that the potential for infringing outputs is irrelevant to the question of training data because fair use must be assessed on a use-by-use basis. The government further contended that “outputs lacking in substantial similarity cannot cause the sort of market harm that is cognizable in the fair-use analysis.” Looking ahead, the government also argued that “any prospective remedy for infringement would need to be limited to what is necessary to afford plaintiffs complete relief.” In the government’s view, allegations concerning particular outputs should not dictate the legal treatment of the distinct activities involved in developing and training LLMs.

Piracy. Notably, the government’s fair-use analysis does not depend on how the works were acquired. It sidesteps the question of whether training on pirated materials is fair use. As a result, the Statement leaves unresolved how courts should treat allegedly pirated training datasets.

Copyright Office. The government’s statement is also notable in that it attacks its own copyright office. It alludes to the removal of the Register of Copyrights, states that her understanding does “not warrant deference” under Loper Bright, and characterizes her reasoning as “threadbare” and ignoring case law.

Takeaways

  • If the courts in the OpenAI litigation and in similar matters adopt the government’s broader view of fair use in the AI-training context, copyright holders should be prepared to take extra measures to protect their works:
    • Contractual frameworks and content distribution. Copyright holders may wish to evaluate whether their contractual framework and current content-distribution model are positioned to preserve the broadest range of rights and remedies. The effectiveness of a given framework may depend on how content is made available, the conditions under which it is accessed, and whether those practices align with the company’s business and enforcement objectives.
    • Alternative statutory claims. Copyright holders may also wish to evaluate whether their content-management practices, including the implementation and maintenance of copyright management information and other content-identification tools, are positioned to support claims under the DMCA or other legal frameworks that may remain available even where traditional copyright infringement theories face additional challenges.
  • The government referenced a GAO report recommending certain legislatively established licensing frameworks and a federal digital-replica protection regime. Copyright holders should be on the lookout for potential legislation.

As courts continue to address the scope of fair use in the AI context, content owners may benefit from assessing whether their contractual frameworks, content-distribution models, content-management practices, and enforcement strategies are positioned to preserve potential claims and maximize available remedies.

© 2026 Sterne, Kessler, Goldstein & Fox PLLC