Contact Us Today! (877) 276-5084

Attorney Steve® Blog

How Retrieval-Augmented Generation (RAG) Is Changing the Legal Landscape

Posted by Steve Vondran | Aug 02, 2026

AI Copyright Law Enters a New Era: Vondran Legal AI trends 2026.

By Attorney Steve® | Vondran Legal®

Introduction

For the past several years, the biggest copyright lawsuits involving artificial intelligence have focused on a single question:

Did AI companies unlawfully copy copyrighted works to train their models?

Authors, publishers, photographers, artists, software developers, and media companies have filed lawsuits against OpenAI, Anthropic, Meta, Stability AI, Midjourney, and others alleging that copyrighted books, images, articles, code, and other creative works were copied without permission during the training process.

But as generative AI technology evolves, so do the legal issues.

Today, a new frontier is emerging that may prove just as significant—if not more so—than the training-data debate.

The focus is shifting from:

  • What went into the model (training), to

  • What comes out of the model (AI-generated outputs).

At the center of this transition is a technology known as Retrieval-Augmented Generation (RAG).

As enterprise AI assistants, legal research tools, AI-powered search engines, customer support bots, and corporate knowledge systems increasingly rely on RAG, courts will likely face entirely new copyright questions.

For businesses developing AI products, understanding these issues is becoming essential.


What Is Retrieval-Augmented Generation (RAG)?

Retrieval-Augmented Generation is an AI architecture that combines two technologies:

  1. A large language model (LLM)

  2. A retrieval engine

Unlike traditional language models that rely only on information learned during training, a RAG system performs an additional step.

Before generating an answer, it searches external sources such as:

  • Websites

  • News articles

  • Company databases

  • Legal documents

  • Technical manuals

  • Scientific papers

  • Internal knowledge bases

  • Licensed content repositories

The retrieved material is then supplied to the language model, which uses it to generate a response.

Think of it as an AI research assistant.

Instead of relying only on memory, it first opens a library, gathers relevant documents, and then writes an answer.


Why RAG Has Become So Popular

Businesses love RAG because it solves one of the biggest limitations of traditional LLMs.

Rather than depending on stale training data, RAG allows AI systems to:

  • Access current information

  • Answer questions about proprietary documents

  • Search internal company knowledge

  • Cite relevant sources

  • Reduce hallucinations

  • Avoid constant retraining

This makes RAG especially attractive for:

  • Law firms

  • Healthcare providers

  • Financial institutions

  • Universities

  • Government agencies

  • Software companies

  • Customer support platforms

  • Enterprise knowledge management

In short, RAG is quickly becoming the default architecture for enterprise AI.


The Copyright Question Has Changed

Early AI copyright lawsuits asked:

Was it lawful to copy copyrighted works to train an AI model?

That remains an enormously important issue.

However, RAG introduces an entirely different legal question.

Instead of focusing on training copies, courts may increasingly ask:

Did the AI reproduce copyrighted expression when responding to a user?

That distinction changes everything.

Rather than investigating datasets collected years ago, courts may examine individual AI responses generated in real time.


Training vs. Inference

One of the most important concepts for businesses to understand is the difference between training and inference.

Training

Training occurs once (or periodically).

Millions—or billions—of documents are processed to teach the model language patterns.

The current lawsuits against OpenAI, Anthropic, Meta, Stability AI, and others largely focus on this stage.


Inference

Inference occurs every time a user asks a question.

With RAG, the AI retrieves external documents and generates an answer based upon those documents.

This is where future copyright litigation may increasingly concentrate.


Why RAG Creates New Copyright Risks

Suppose a user asks:

"Summarize today's CNN article about the Federal Reserve."

A RAG system might retrieve the article and produce an answer.

But what if that answer:

  • reproduces multiple paragraphs,

  • copies the structure,

  • includes expressive wording,

  • or gives users everything they need without visiting CNN?

Now the copyright analysis begins to resemble traditional infringement law.

The issue is no longer merely whether the AI learned from the article.

The issue becomes whether the AI reproduced protected expression.


CNN v. Perplexity AI: A Glimpse Into the Future

One of the most closely watched AI copyright cases was filed by CNN in 2026 against Perplexity AI.

According to the complaint, CNN alleges that Perplexity:

  • copied CNN journalism,

  • reproduced articles,

  • displayed photographs,

  • incorporated videos,

  • and generated AI responses that closely tracked CNN reporting.

CNN contends that users can obtain the substance of its journalism without ever visiting CNN's website.

Importantly, the lawsuit focuses not merely on model training but on how Perplexity allegedly retrieves and presents copyrighted content to users.

As of this writing, these allegations remain contested, and no court has ruled on the merits.

Nevertheless, the case illustrates where AI copyright litigation may be headed.


Other Publishers Have Raised Similar Concerns

CNN is hardly alone.

Other publishers have asserted comparable claims.

These include:

  • Dow Jones

  • The New York Post

  • BBC

These organizations argue that AI-powered search engines may:

  • summarize copyrighted journalism,

  • reduce website traffic,

  • replace original reporting,

  • diminish advertising revenue,

  • undermine subscription models.

Whether courts ultimately agree remains to be seen.


Traditional Copyright Principles Still Apply

Although AI is new, copyright law is not.

Most future disputes will likely be analyzed using familiar doctrines.

Courts may ask:

1. Was Protected Expression Copied?

Copyright protects expression—not facts.

AI may freely discuss historical facts.

It cannot freely reproduce creative expression.


2. Is the Output Substantially Similar?

Courts often compare:

  • wording,

  • structure,

  • organization,

  • sequence,

  • overall expression.

The closer an AI output comes to the original work, the greater the legal risk.


3. How Much Was Copied?

Small quotations may be permissible.

Large excerpts present greater concerns.

Entire paragraphs copied verbatim are generally much more problematic than concise summaries.


4. Is the Use Transformative?

Fair use often turns on whether the new use adds something genuinely different.

A transformative AI response may:

  • synthesize multiple sources,

  • compare viewpoints,

  • provide commentary,

  • analyze competing opinions.

Simply restating an article in different words may receive less favorable treatment.


5. Does the Output Replace the Original Work?

Perhaps the biggest question.

If users no longer need to visit the publisher's website because the AI provides essentially everything, courts may consider whether the AI acts as a market substitute.

Market substitution has long been an important factor in fair-use analysis.


Important Copyright Cases That May Shape This Area

Although RAG-specific precedent is still developing, several landmark Supreme Court decisions are likely to influence future litigation.

Harper & Row Publishers, Inc. v. Nation Enterprises, 471 U.S. 539 (1985)

The Supreme Court emphasized that unauthorized publication of the "heart" of a copyrighted work can weigh heavily against fair use, particularly where the use harms the copyright owner's market.

https://supreme.justia.com/cases/federal/us/471/539/


Campbell v. Acuff-Rose Music, Inc., 510 U.S. 569 (1994)

This landmark case established the modern understanding of transformative use and remains central to virtually every fair-use analysis.

https://supreme.justia.com/cases/federal/us/510/569/


Google LLC v. Oracle America, Inc., 593 U.S. ___ (2021)

The Court held that Google's limited copying of portions of the Java API constituted fair use under the specific facts presented, emphasizing the highly functional nature of computer interfaces and the transformative purpose of the use. While not a RAG case, it demonstrates how technological innovation can influence fair-use analysis.

https://supreme.justia.com/cases/federal/us/593/19-956/


Andy Warhol Foundation v. Goldsmith, 598 U.S. 508 (2023)

The Supreme Court narrowed the concept of transformative use, emphasizing that courts must examine whether the secondary use serves a substantially different commercial purpose.

This decision will likely play an important role in future AI litigation.

https://supreme.justia.com/cases/federal/us/598/21-869/


Enterprise AI Developers Face Unique Risks

Many organizations mistakenly believe that copyright concerns apply only to public AI search engines.

That is incorrect.

Enterprise RAG systems often retrieve:

  • employee manuals,

  • engineering specifications,

  • software documentation,

  • research papers,

  • customer contracts,

  • licensed databases,

  • subscription services.

Without proper safeguards, AI assistants may inadvertently disclose or reproduce protected material beyond its authorized audience.


Best Practices for Organizations Deploying RAG

Businesses should consider implementing robust governance policies, including:

Limit Verbatim Copying

Favor summaries over wholesale reproduction.


Use Short Quotations

Quote only what is reasonably necessary.


Attribute Sources

Whenever practical, identify the originating publication or document and provide links to the original source.


Respect Access Controls

Do not allow AI to expose documents users are not authorized to view.


Maintain Retrieval Logs

Logging can assist with:

  • compliance,

  • audits,

  • troubleshooting,

  • litigation holds,

  • internal investigations.


Establish Output Policies

Organizations should adopt rules governing:

  • maximum quotation length,

  • citation requirements,

  • restricted content,

  • confidential information,

  • licensed materials.


Review Licensing Agreements

Many commercial databases prohibit automated extraction or AI-based redistribution.

Contract claims may accompany copyright claims.


Train Employees

Employees should understand that AI-generated responses may still create intellectual property risks.

AI governance is increasingly becoming part of corporate compliance programs.


Key Takeaways

The next chapter of AI copyright law is already unfolding.

While training-data lawsuits remain active, courts are beginning to confront a different question:

What happens when AI retrieves copyrighted works in real time and reproduces portions of those works for users?

For businesses building AI products, the implications are significant:

  • AI copyright compliance is no longer limited to training datasets.

  • Retrieval-Augmented Generation introduces new risks involving AI outputs.

  • Publishers are increasingly challenging AI systems that summarize or reproduce journalism.

  • Traditional copyright principles—including substantial similarity, fair use, market substitution, and transformative use—are likely to remain the legal framework courts apply.

  • Organizations deploying RAG should implement technical, contractual, and governance safeguards to minimize infringement risk.

  • AI developers should pay close attention to emerging cases involving Perplexity AI and other retrieval-based platforms, as these decisions may shape the next generation of AI copyright law.

Final Thoughts

Artificial intelligence is rapidly changing how information is created, accessed, and consumed. As Retrieval-Augmented Generation becomes the foundation of enterprise AI, legal scrutiny will increasingly shift from how models are trained to how AI systems retrieve, synthesize, and present copyrighted content at the moment a user asks a question.

Companies that proactively implement AI governance, copyright compliance procedures, licensing reviews, and output controls will be far better positioned to innovate while reducing legal risk. For developers, publishers, and businesses alike, understanding the distinction between training and inference may become one of the defining legal issues of the AI era.

About the Author

Steve Vondran
Steve Vondran

Thank you for viewing our blogs, videos and podcasts. As noted, all information on this website is Attorney Advertising. Decisions to hire an attorney should never be based on advertising alone. Any past results discussed herein do not guarantee or predict any future results. All blogs are written by Steve Vondran, Esq. unless otherwise indicated. Our firm handles a wide variety of intellectual property and entertainment law cases from music and video law, Youtube disputes, DMCA litigation, copyright infringement cases involving software licensing disputes (ex. BSA, SIIA, Siemens, Autodesk, Vero, CNC, VB Conversion and others), torrent internet file-sharing (Strike 3 and Malibu Media), California right of publicity, TV Signal Piracy, and many other types of IP, piracy, technology, and social media disputes. Call us at (877) 276-5084. AZ Bar Lic. #025911 CA. Bar Lic. #232337

Contact us for an initial consultation!

For more information, or to discuss your case or our experience and qualifications please contact us at (877) 276-5084. Please note that our firm does not represent you unless and until a written retainer agreement is signed, and any applicable legal fees are paid. All initial conversations are general in nature. Free consultations are limited to time and availability of counsel and will depend on the type of case you are calling about (no free consultations for other lawyers). All users and potential clients are bound by our Terms of Use Policies. We look forward to working with you!
The Law Offices of Steven C. Vondran, P.C. BBB Business Review

Menu