Topics · 01

Training Data & Intellectual Property in Generated Content

Verify legal bases for training, fine-tuning and RAG ingestion, preserve human creative input, and manage commercialization, licensing and infringement claims.

How can model training data be legally verified, and can AI-generated content be safely commercialized?

Key Trigger Scenarios

  • Procuring or preparing datasets for pre-training and fine-tuning
  • Ingesting client materials, proprietary data or scraped content into RAG systems
  • Commercialising and licensing generated text, images, audio, video or code
  • Receiving copyright infringement notices or detecting substantial similarity risks

Preliminary Due Diligence & Materials

  1. 01Map data sources, commercial licences, and fair use / text-and-data mining exceptions
  2. 02Preserve prompts, model parameters, iterative revisions, and human contribution logs
  3. 03Audit model licences, open-source conditions, and upstream data usage policies
  4. 04Establish rapid-response takedown, counter-notice, and evidence preservation protocols

Selected Cases

Judicial Trends & Regulatory Standards

Read each case in its procedural context. Related cases may offer comparisons across technologies; they do not establish a single rule for every system.

China

Beijing AI Text-to-Image Case

Evaluating prompt design, parameter adjustments, and the generation process, the court held that the disputed image reflected the user's intellectual investment and personalized expression, determining copyrightability and ownership on this basis.

First-instance judgment

  • The court held that the disputed image satisfied the requirement of originality.
  • Artificial intelligence models lack legal personality, making human input during the specific generation process a key basis for determining authorship.
View Case Study →
China

Guangzhou Ultraman AI Image Case

The defendant connected to a third-party AI image service through an API. Users could generate images substantially similar to the protected Ultraman works. The court classified the defendant as a generative AI service provider and found infringement of the reproduction and adaptation rights.

First-instance judgment

  • The court found that some outputs reproduced original elements of Ultraman and that others retained those elements while adding new features, infringing the reproduction and adaptation rights. It did not separately repeat the analysis of the information-network dissemination right.
  • Because the defendant provided the AI image service to users through an API, the court treated it as a generative AI service provider and required technical measures preventing substantially similar outputs during ordinary use of Ultraman-related prompts.
View Case Study →
United Kingdom

UK Getty Images Case

Getty Images alleged that Stable Diffusion infringed intellectual property rights through its training, distribution, and outputs bearing Getty Images or iStock watermarks. Before the close of trial, Getty Images abandoned its training, output-copyright, and database-right claims. The court dismissed secondary copyright infringement and found only limited historical instances of trademark infringement.

First-Instance Judgment; partially under appeal

  • The court held that an intangible model may be an article, but Stable Diffusion was not an infringing copy because its model weights had never stored or reproduced the copyright works; the secondary infringement claim was dismissed.
  • Getty Images succeeded only in relation to limited examples of iStock and Getty Images watermarks generated by certain earlier model versions, and the court stressed the narrow and historic scope of those findings.
View Case Study →
United States

U.S. ROSS Training Data Case

ROSS used training materials derived from Westlaw headnotes to develop a legal search tool. The district court found numerous headnotes protectable and rejected the fair use defense. The Third Circuit has accepted an interlocutory appeal.

Partial Summary Judgment; Appeal Pending

  • The district court granted partial summary judgment in favor of Thomson Reuters regarding 2,243 headnotes.
  • The district court determined that the training use was competitive and rejected the fair use defense.
View Case Study →
United States

U.S. Anthropic Book Training Case

The court distinguished model training and the digitization of purchased print books from downloads from pirated repositories before granting final approval to a $1.5 billion class action settlement. The ruling evaluated training activities separately from data acquisition methods.

Class Action Settlement Granted Final Approval

  • The court held that model training and the digitization of lawfully purchased print books constituted fair use on the record presented.
  • The court held that obtaining and storing books from pirated repositories required independent scrutiny that downstream training purposes could not automatically excuse.
View Case Study →
United States

U.S. Meta Book Training Case

Authors alleged that Meta used copyrighted books without authorization to train its Llama models. Based on the evidentiary record, the court granted summary judgment in favor of Meta against certain plaintiffs, emphasizing that market harm evidence remains central to outcomes in related cases.

Partial Summary Judgment Granted; Proceedings Ongoing

  • The court determined that Meta's use of the books for model training was transformative.
  • Because the plaintiffs failed to submit sufficient evidence of relevant market harm, Meta obtained summary judgment on those plaintiffs' copyright claims.
View Case Study →

Related Services

Tailored Legal Services for This Scenario

Regulatory Frameworks

Applicable Regulatory Frameworks

Research team

AI Legal Research Center

Consult on This Topic →