The DMCA may no longer provide a shortcut some copyright owners hoped for against generative AI outputs. The prospect of statutory damages, no fair-use fight, and no need to prove copyright infringement made § 1202 claims particularly appealing. But in Doe v. GitHub, the Ninth Circuit rejected the plaintiff’s output theory, holding that, as alleged, Copilot generates new works through statistical prediction rather than stripping copyright management information (CMI) from existing copies. The decision also provides important guidance on the “identicality requirement,” explaining that identicality is not an independent element of a DMCA claim but may serve as evidence that CMI was removed from a copy of a copyrighted work. The ruling narrows the circumstances in which AI outputs can support §1202 claims, potentially pushing copyright owners toward training-stage CMI theories and traditional copyright infringement claims based on AI outputs.
The Parties and the Dispute
Plaintiffs are anonymous programmers who published open-source code on GitHub under licenses requiring attribution. Defendants are GitHub, its parent Microsoft, and the OpenAI entities.
The accused products are GitHub, Copilot, and OpenAI Codex, large language models trained on public GitHub repositories. As described by the appeals court, “GitHub is a platform that allows developers to store, manage, and share their code,” and “GitHub, Copilot, and Codex are tools that help programmers write code using artificial intelligence—specifically, large language models that were trained on millions of software projects on GitHub.”
Plaintiffs presented two theories of DMCA liability: (1) an “input” theory “that defendants violated section 1202(b)(1) at the training stage by removing CMI from putative class members’ code before feeding the stripped code into Copilot as training data,” and (2) an “output” theory “that Copilot sometimes returns memorized training data as output to users without including the source code’s CMI, and that in doing so Copilot ‘remove[s]’ or ‘alter[s]’ CMI on copies of plaintiffs’ protected work, in violation of section 1202(b)(1) and (3).” Defendants argued that plaintiffs lacked standing and that § 1202(b) reaches only removal and alteration of CMI from a copy of the work, a concept the parties and the appeals court referred to as an “identicality” requirement. The district court dismissed because the alleged outputs were not sufficiently identical to the copyrighted code to support a CMI-removal claim.
The district court certified for interlocutory appeal the question “whether Sections 1202(b)(1) and (b)(3) of the DMCA impose an identicality requirement.”
The Ninth Circuit’s Reasoning
Standing. The appeals court found a plausible “substantial risk” of injury. The complaint cited research on memorization in large models, GitHub’s own duplicate-detection filter for 150-character matches, and examples of named plaintiffs’ code appearing verbatim. That sufficed at the pleading stage.
Forfeiture of the input theory. Plaintiffs’ counsel conceded that copying training data “perhaps” did not violate the licenses. The district court stated the “complaint is not about training.” Plaintiffs never objected. The Ninth Circuit declined to consider the theory.
Merits of the output theory. Based on plaintiffs’ own description of Copilot, the appeals court concluded that the AI tool generates new works through a probabilistic process rather than making copies of existing works. That “cannot reasonably be described as a copy of” the protected code. Starting with text, the appeals court read “remove” and “alter” to require an affirmative act on CMI attached to “a work that already exists”—i.e., an existing copy. Failing to include CMI in a new work, even a similar or derivative work, does not constitute removal or alteration because no CMI was attached to that work in the first place.
While the district court and defendants characterize this principle as the DMCA’s “identicality” requirement, the appeals court rejected “identicality” as a freestanding element. The appeals court explained that “identicality” is best understood as a gloss on the statutory terms “remove,” “alter,” and “copies” rather than an independent element of a § 1202(b) claim. Instead, identicality is evidence. Near-identical reproduction without CMI supports an inference of removal. Material differences suggest a new work to which CMI was never attached. The defendants conceded that “two works need not be literally identical to support an inference that CMI has been removed.”
The decisive point was plaintiffs’ own description of Copilot. The complaint alleged a “complex probabilistic process” that infers statistical patterns and predicts the most likely completion. That describes generation, not retrieval. The appeals court contrasted Copilot with a search engine, which retrieves stored copies. Had plaintiffs alleged that Copilot returned stored copies without CMI, they “might have a stronger claim.”
The appeals court also warned against letting § 1202(b) swallow infringement law. Substantial similarity without attribution describes an ordinary copyright case. Treating it as a DMCA violation would expose defendants to up to $25,000 per violation under § 1203(c)(3), with no per-work cap like the $30,000 ceiling in § 504(c)(1). The appeals court called that exposure “potentially ruinous.”
What This Means for AI Litigation
First, the key takeaway for litigants is the importance of input theories for DMCA claims. The training-stage theory appears more consistent with the appeals court’s reading of § 1202(b)(1) because it focuses on alleged acts taken with respect to existing copies containing CMI. The Ninth Circuit did not reach that theory, however, concluding that plaintiffs had forfeited it. Future plaintiffs asserting DMCA claims may therefore seek to plead and preserve training-stage removal theories more explicitly.
Second, output claims may be best pursued under the Copyright Act, not the DMCA. The appeals court left copyright infringement open. Plaintiffs relying on memorized output may increasingly look to traditional infringement theories, where substantial similarity, copying, and fair-use defenses are likely to remain central issues.
Third, pleading of the technology can be outcome-determinative. The appeals court held plaintiffs to their own account of how Copilot works. Plaintiffs who can plausibly allege that a system retrieves or reproduces stored works, rather than generating outputs through statistical prediction, may face a different § 1202 analysis.
Provenance and the Evidentiary Gap
Plaintiffs’ claims failed partly for want of a traceable act. Current models do not retain provenance in any form a litigant can inspect. That could change. If developers build systems that carry machine-readable attribution, copyright notices, ownership information, license terms, or content credentials (e.g., the C2PA standard) through training and into output, the record would show where CMI existed and where it disappeared. Much of that information falls within the categories of CMI identified in § 1202(c).
More comprehensive provenance systems could benefit both sides. Plaintiffs may view such records as evidence concerning the presence or removal of attribution information, while developers may view them as evidence of authorization, licensing, or compliance.
The EU Angle
The EU AI Act, Regulation (EU) 2024/1689, Article 53(1), already requires general-purpose model providers to keep technical documentation on training data and to publish a “sufficiently detailed summary about the content used for training.” They do not yet require item-level provenance. But if the regime evolves toward granular tracking, it may build the record U.S. plaintiffs need. To the extent AI Act compliance documentation contains information about training data practices, litigants may seek such materials in future U.S. litigation, subject to ordinary discovery rules and applicable limitations.
Takeaway
Section 1202(b) is not a general attribution right. It reaches acts on existing copies. In the Ninth Circuit, output-based DMCA claims against generative models are likely difficult to sustain absent allegations of retrieval-type behavior or removal of CMI from existing copies. The action shifts to training-stage removal and infringement. Provenance infrastructure, wherever it originates, will decide how far that shift goes.
Related Industries
Related Services
Receive insights from the most respected practitioners of IP law, straight to your inbox.
Subscribe for Updates