Briefly
Case Law

Nvidia Sued Over Alleged Piracy of Books3 Dataset for AI Training

United States·Courthouse News Service·⏱️ 3 min readBriefly Analysis

Summary

  • Nvidia faces an avalanche of lawsuits over its AI model training program due to alleged use of pirated materials.
  • The company's large language models were trained on datasets containing copyrighted works without permission.
  • Investors blast Nvidia for using stolen materials, leading to a significant decline in the company's stock price.

What Happened

Nvidia is facing an avalanche of lawsuits over its AI model training program, with investors blasting the company for using pirated materials in its development process. The suits claim that Nvidia's large language models were trained on datasets containing copyrighted works without permission, including full-length books, YouTube videos, and thousands of hours of human speech recordings. One stockholder has filed a suit against Nvidia's board of directors and CEO Jensen Huang, alleging that they knowingly allowed the company to use stolen materials for financial gain. The plaintiff claims that the defendants failed to exercise control over the illegal acts, despite approving proxy statements that touted the company's focus on information security and privacy protections. The suits have led to a significant decline in Nvidia's stock price, from around $624 in January 2024 to $1,208 in June 2024.

Relevant Legal/Regulatory Context

The lawsuits against Nvidia stem from the use of a dataset known as 'The Pile,' which included a subcollection of nearly 200,000 pirated books called Books3. The plaintiffs claim that Nvidia used The Pile to train models in its Megatron line without obtaining necessary licenses and permissions. This is not an isolated incident, as similar suits have been filed against Nvidia and Microsoft for allegedly extracting unique biometric signatures from recordings of individuals' voices. The cases highlight the importance of ensuring that data used in AI development is properly licensed and permissioned to avoid significant reputational damages and liability.

Why It Matters

The Nvidia case serves as a warning to companies using pirated materials in their AI model training programs. The use of such materials can lead to significant reputational damages and liability, as seen in the decline of Nvidia's stock price following the disclosure of the lawsuits. Lawyers and compliance officers should be aware of the potential risks associated with using unlicensed data in AI development and take steps to ensure that all necessary permissions are obtained. This is particularly important given the increasing use of AI models in various industries, where the consequences of copyright infringement can be severe.

Practical Implications

Lawyers and compliance officers should be aware of the potential for companies to use pirated materials in their AI model training programs, which can lead to significant reputational damages and liability. This case highlights the importance of ensuring that data used in AI development is properly licensed and permissioned.

Source

Source: Original reporting via Courthouse News Service

AI Business Impact

How does this affect your business?

Get an AI analysis of this article grounded in your jurisdictions, practice areas, and any policy documents you've uploaded to Wansom.

Wansom is AI and can make mistakes.