AI Copyright Litigation Enters a New Phase: Apple's "Access vs. Use" Defense Could Reshape Generative AI Law
By Attorney Steve® | Vondran Legal®
The first wave of artificial intelligence copyright litigation focused on a relatively straightforward question:
Can AI companies lawfully train their models using copyrighted works without obtaining permission from the copyright owner?
That question remains far from settled.
However, a new legal battle involving Apple may shift the conversation toward an even more fundamental issue:
When does merely accessing publicly available internet content become unlawful use under copyright law and the Digital Millennium Copyright Act (DMCA)?
If Apple's arguments gain traction, this case could become one of the most influential AI copyright decisions to date, affecting not only Apple, but virtually every company developing large language models (LLMs), image generators, search engines, and other AI technologies.
The Lawsuit Against Apple
Apple is defending a proposed class action pending in the United States District Court for the Northern District of California involving allegations that it scraped millions of publicly available YouTube videos for AI training.
According to the complaint, Apple allegedly:
- collected YouTube videos,
- copied copyrighted audiovisual works,
- bypassed YouTube's technological restrictions,
- and used those videos to develop internal AI foundation models.
Rather than arguing primarily about fair use, Apple reportedly focuses on a different legal theory:
The videos were publicly available to everyone.
According to Apple's motion to dismiss, no passwords, subscriptions, or restricted credentials were required to watch the videos. Therefore, Apple argues it did not circumvent technological protection measures prohibited under the DMCA.
That distinction may ultimately become more important than the traditional fair-use analysis.
The New Question: Access vs. Use
Historically, copyright cases ask questions like:
- Was the work copied?
- Was the copying authorized?
- Was the use fair?
Apple's defense introduces an earlier question:
What legally counts as "access"?
Simply because information is copyrighted does not necessarily mean viewing it constitutes unlawful access.
The internet is filled with copyrighted material that anyone may lawfully view.
Examples include:
- newspaper articles
- YouTube videos
- blog posts
- photographs
- software documentation
- public GitHub repositories
The legal question becomes:
Can an AI system lawfully collect publicly accessible information without violating the DMCA?
That issue is surprisingly unsettled.
Why the DMCA Changes the Analysis
Many AI lawsuits primarily involve the Copyright Act.
Apple's case raises issues under the Digital Millennium Copyright Act, particularly:
- 17 U.S.C. §1201
- 17 U.S.C. §1202
These sections serve different purposes.
Section 1201
Section 1201 prohibits:
- circumventing technological protection measures
- bypassing access controls
- defeating encryption or digital locks protecting copyrighted works
Apple argues that publicly viewable YouTube videos were not protected by qualifying technological access controls because anyone could watch them without authorization credentials.
If accepted, that argument could significantly narrow future DMCA claims involving publicly accessible websites.
Section 1202
Section 1202 addresses something entirely different.
Rather than focusing on access, it protects Copyright Management Information (CMI).
CMI includes information such as:
- author names
- copyright notices
- ownership information
- licensing terms
- attribution data
Many AI training pipelines clean enormous datasets before model training.
During that cleaning process:
- metadata may be stripped,
- filenames removed,
- attribution discarded,
- copyright notices deleted.
Whether those actions violate Section 1202 has become one of the hottest issues in AI litigation.
Recent Cases Defining the DMCA Landscape
Several important federal decisions illustrate how courts are approaching these claims.
1. Raw Story Media v. OpenAI
In one of the earliest DMCA-only AI lawsuits, publishers alleged OpenAI removed copyright management information from news articles during AI training.
The Southern District of New York dismissed the complaint after concluding the plaintiffs lacked Article III standing because they failed to demonstrate a sufficiently concrete injury tied to the alleged removal of CMI. Later efforts to revive the claims through reconsideration were also unsuccessful.
Case Link:
https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1%3A2024cv01514/
Practical takeaway:
Merely alleging removal of metadata may not be enough. Plaintiffs often need to connect that removal to a concrete injury recognized by the court.
2. The New York Times Co. v. Microsoft & OpenAI
Perhaps the highest-profile AI copyright case currently pending.
Unlike Raw Story, portions of the DMCA claims survived dismissal.
Judge Sidney Stein allowed certain Section 1202 claims to proceed while dismissing others without prejudice, illustrating that DMCA claims are highly fact-dependent rather than categorically barred.
Case Link:
https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2023cv11195/
The decision demonstrates:
- some DMCA theories remain viable;
- others require more specific factual allegations;
- courts are carefully distinguishing between different subsections of §1202.
3. Intercept Media Litigation
The Intercept similarly asserted that OpenAI removed copyright management information while training its AI models.
Unlike Raw Story, the complaint included significantly more detailed factual allegations and examples, helping the claims survive the pleading stage according to reports covering the litigation.
This reinforces an emerging trend:
Specific factual allegations matter.
Why Apple's Defense Could Become a Landmark
If Apple prevails, the decision could establish an important principle:
Public accessibility alone does not create DMCA liability.
That would have enormous implications.
Every major AI developer relies, at least in part, upon publicly available internet information.
Potential beneficiaries could include:
- OpenAI
- Anthropic
- Google DeepMind
- Meta
- Amazon
- Apple
- Microsoft
- xAI
- future AI startups
The case may influence how courts distinguish:
- public webpages
- password-protected content
- licensed databases
- subscriber-only materials
- APIs
- robots.txt restrictions
- technical anti-scraping measures
What About Copyright Infringement?
Even if Apple succeeds on its DMCA arguments, copyright claims do not disappear.
Instead, litigation shifts toward different questions.
For example:
Did Apple make unauthorized copies?
Were those copies transformative?
Does AI training qualify as fair use?
Were outputs substantially similar?
Did memorization occur?
Those issues remain pending across numerous AI cases.
Practical Compliance Lessons for AI Companies
Regardless of how Apple's motion is decided, companies developing AI systems should already be implementing defensible compliance procedures.
1. Document Data Sources
Know exactly where training data originated.
Maintain records regarding:
- licenses
- datasets
- web crawls
- public repositories
- API permissions
2. Preserve Copyright Metadata
Avoid automatically deleting:
- copyright notices
- author names
- licensing terms
- attribution information
Metadata preservation may become increasingly important under Section 1202.
3. Differentiate Public Content from Restricted Content
Not all internet content is equally accessible.
Developers should distinguish between:
- publicly viewable pages;
- password-protected content;
- subscriber-only materials;
- contractual restrictions;
- technical access controls.
4. Maintain Written AI Governance Policies
Modern AI companies should establish written policies addressing:
- data collection;
- web scraping;
- copyright compliance;
- licensing;
- dataset auditing;
- takedown procedures;
- metadata preservation.
These policies may become valuable evidence in future litigation.
5. Evaluate Licensing Where Appropriate
Even if some publicly available data may be lawfully accessible, licensing can significantly reduce litigation risk, especially for:
- publishers;
- stock photography;
- music;
- software documentation;
- proprietary databases.
Looking Ahead
The first generation of AI lawsuits largely asked:
Can copyrighted works train AI models?
The next generation appears ready to ask something even more fundamental:
What legally constitutes "access" to information on the public internet?
That distinction may ultimately shape:
- AI copyright law,
- DMCA jurisprudence,
- web scraping practices,
- enterprise compliance programs,
- licensing negotiations,
- and the architecture of future AI training pipelines.
Apple's motion may therefore become one of the most closely watched copyright decisions of the AI era. While no one can predict how the court will rule, the case underscores a broader reality: successful AI development increasingly depends not only on technical innovation, but also on legally defensible data governance.
Need Help With AI Copyright or DMCA Issues?
Vondran Legal® represents clients nationwide in intellectual property, AI, software, copyright, DMCA, and technology disputes. If your company is facing questions involving AI training data, web scraping, copyright compliance, or Digital Millennium Copyright Act issues, experienced legal counsel can help evaluate your risk and develop practical compliance strategies before disputes arise.
Disclaimer: This article is for educational purposes only and does not constitute legal advice. Every case depends on its specific facts, applicable law, and the jurisdiction in which it arises.

