What AI Intellectual Property Theft Means for Students
Students hear that AI 'steals' work but rarely get a clear explanation of what that means legally and technically. This guide breaks down the two core legal questions behind AI copyright lawsuits and what the 2025 rulings actually say.
- SAT
- ACT
- GRE
- MCAT
- ASVAB
- digital-sat
- adaptive-testing
- registration-fee
- content-outline
- score-target
A student sees the phrase “AI steals art” in a post, then reads that a court allowed some AI training under fair use, then hears that another court rejected fair use in a different AI case. At that point, “steals” has stopped being an explanation. It has become a bundle of accusations that need to be separated before anyone can argue about them responsibly.
As of Q3 2026, the legal picture is active and mixed. In 2025 alone, U.S. courts and the U.S. Copyright Office pointed in different directions depending on the facts: one ruling rejected fair use for AI training in a direct-competition setting, a Copyright Office report warned that some training uses go beyond established fair-use boundaries, and another court treated legally purchased books differently from unlicensed copies in an AI training dispute.[1][2][3]
For students, the cleanest starting point is this: arguments about AI intellectual property theft usually involve two different questions.
| Question | Plain-English version | What the law is trying to decide |
|---|---|---|
| Input-side copying | What copyrighted works went into the AI system during training? | Whether copying works for training without permission is infringement or fair use. |
| Output-side infringement | What does the AI system produce for users? | Whether generated text, images, music, or other outputs are too close to protected works. |
Those two questions can overlap, but they are not the same. A company might argue that its training process is lawful even if a particular generated output later infringes someone’s copyright. A creator might argue that both the training copies and the outputs are unlawful. If a debate does not say which side it means, it is probably moving too fast.

What “training on” a work actually means
When people say an AI model was “trained on” books, images, songs, code, or articles, they usually mean that copies of those works were used as data while the system learned statistical patterns. The model is not normally flipping through a book like a student in a library. Training involves computational processing of large collections of material so the system can learn relationships among words, pixels, sounds, or other features.
That is why the popular analogy “AI learns like a human” does not settle the copyright issue. The U.S. Copyright Office’s 2025 Part 3 report rejected that analogy, distinguishing human memory’s “imperfect impressions” from AI training processes that can involve making “perfect copies” of works.[2]
This does not automatically mean every training copy is illegal. It means the legal question is not answered by saying “learning is allowed.” Copyright law asks more specific questions: What was copied? Why was it copied? Was the copying necessary for that use? Is the new use transformative? Does it compete with the original market or a market the copyright owner could reasonably license? Those questions are where the 2025 cases become useful.
The input-side question: is training fair use?
Fair use is the part of U.S. copyright law that can allow some unauthorized uses of copyrighted material. It is not a magic label. Courts look at factors such as the purpose and character of the use, the nature of the copyrighted work, the amount used, and the effect on the market for the original.
With AI training, companies often argue that copying is transformative because the system is not republishing the original works in the ordinary way. Copyright owners answer that the copying can be massive, commercial, and aimed at building products that compete with the very people whose work supplied the training material.
The 2025 developments matter because they show that courts are not treating “AI training” as one single category.
Thomson Reuters v. Ross: direct competition mattered
In February 2025, a federal court ruled against Ross Intelligence’s fair-use defense after Ross used copyrighted Westlaw headnotes from Thomson Reuters to train a legal research tool. Judge Stephanos Bibas found that the use was not fair use, emphasizing that Ross and Thomson Reuters were direct competitors and that the copying was “not reasonably necessary.”[1]
That detail is not decorative. It changes the legal feel of the case. A tool trained on a competitor’s legal headnotes to build a competing legal research product looks different from a research project, a parody, or a classroom excerpt. Market harm is one of the fair-use factors, and this case put competition close to the center.
There is also a caution: Ross was not a large language model like ChatGPT. It involved a non-generative AI legal search tool. So the case is important, but it should not be flattened into “all AI training is illegal.” A better student claim would be narrower: in a 2025 case involving a legal AI tool and a direct competitor’s headnotes, a court rejected fair use for training-related copying.[1]
The Copyright Office report: training can cross fair-use boundaries
In May 2025, the U.S. Copyright Office released Part 3 of its Copyright and Artificial Intelligence report. The report concluded that using copyrighted works to train AI systems that generate competing expressive content can go “beyond established fair use boundaries.”[2]
That is a strong institutional statement, but it is not the same thing as a Supreme Court ruling. The report gives the Copyright Office’s position on how copyright principles should apply; it does not end the lawsuits. It is also worth noting that the Part 3 report was released as a pre-publication version under unusual political circumstances, so students should cite it as an important government position, not as final nationwide law.[2][4]
The report is especially useful for one common debate move. If someone says, “AI training is just like a person reading a book,” the Copyright Office gives a reason to slow down: human learning and machine training are not identical in how copying occurs, how outputs can be generated at scale, or how the resulting system may compete in expressive markets.[2]
Anthropic: purchased books and unlicensed datasets were treated differently
The Anthropic rulings in July 2025 are probably the clearest example of why “AI training is legal” and “AI training is theft” are both too blunt. Judge William Alsup found that using legally purchased books for large language model training was fair use. But he treated unlicensed works from the Pile dataset differently, finding that Anthropic’s use of those unlicensed works was not protected in the same way.[3]
That split matters because it separates the training purpose from the acquisition trail. A company may argue that training itself is transformative, but a court may still care intensely about whether the works were lawfully obtained, whether copies were retained, and whether unauthorized libraries or datasets were used along the way.
Anthropic also proposed a $1.5 billion settlement to roughly 500,000 authors, which would have been the largest copyright recovery in U.S. history, but the court rejected that proposal.[3] The settlement detail shows the scale of the dispute; the rejection is the part that keeps the detail honest.
| 2025 development | What happened | What students should not overclaim |
|---|---|---|
| Thomson Reuters v. Ross | Court rejected fair use where a legal AI company used copyrighted legal headnotes while competing with the copyright owner. | It did not decide every generative AI training case. |
| U.S. Copyright Office Part 3 report | Office said some training for systems that generate competing expressive content goes beyond established fair-use boundaries. | It is an influential government position, not a final court ruling. |
| Anthropic rulings | Court treated legally purchased books as fair-use training material but treated unlicensed works from the Pile differently. | It did not make either side’s broad slogan true. |
The output-side question: did the AI generate something infringing?
Output-side infringement is easier to picture. If an AI system generates an image that is substantially similar to a protected illustration, a song that copies protected lyrics, or text that closely tracks a copyrighted article, the legal argument is about the product that came out, not only the data that went in.
This is where students need to separate style, idea, and expression. Copyright generally protects particular expression, not a broad idea or genre. “A gloomy fantasy castle” is not the same kind of claim as a near-copy of a specific painting. “A pop song about heartbreak” is not the same as copying protected lyrics. The closer the output gets to protected expression, the more serious the output-side issue becomes.
Many public arguments mix the two sides. A plaintiff may say the model was trained on copyrighted material without permission and that the system can produce infringing outputs. A company may deny both, or it may argue that training is lawful while also trying to reduce outputs that imitate specific works. Those are different defenses to different claims.
Why so many people are suing
The lawsuits are not limited to one art form or one company. Reporting and litigation trackers have identified more than 40 active lawsuits involving parties such as The New York Times, Getty Images, major book publishers, Disney, music labels, individual authors, and visual artists.[5]
That number is useful for scale, not for panic. A high lawsuit count means the issue is broad and economically important. It does not tell you, by itself, who is right. Lawsuits contain allegations; rulings tell us what a court has actually accepted, rejected, or allowed to continue.
The dispute is also not only American. In November 2025, a German court ruled in GEMA v. OpenAI that OpenAI violated copyright law by training on song lyrics, showing that courts outside the United States are also being asked to decide how copyright applies to AI training and outputs.[5]
Why this feels different from ordinary plagiarism
Students often meet this issue through school plagiarism rules, so it is tempting to treat AI copyright lawsuits like a giant academic honesty case. That comparison helps only up to a point.
If a student submits AI-generated text as their own work without disclosure, the school’s concern is usually authorship, learning, and honesty. The question is whether the student represented someone—or something else’s—work as their own. A 2024 ICAI survey cited by Krater reported that 43% of students admitted using AI tools and 18% submitted AI-generated content as their own without disclosure.[6]
Copyright litigation asks a different set of questions. Did a company copy protected works? Was that copying licensed or fair use? Did the system generate infringing outputs? Did the use harm a real or potential market? A student can violate a school AI policy without committing copyright infringement, and a company can face copyright claims even if no student ever turns in the output.
There is still a tension students notice for good reason. Schools often give students clear rules about disclosure and consequences, while companies have spent years building AI systems in a legal environment that is still being sorted out. That does not prove the companies are guilty in every case. It does explain why “just follow the rules” can sound thin when the rules for large-scale training are still contested.
How to evaluate the claim “AI steals IP”
A useful response is not to accept or reject the phrase immediately. Ask what kind of claim is being made.
- Is the claim about inputs—copyrighted books, images, songs, code, or articles used during training?
- Is the claim about outputs—AI-generated material that allegedly copies protected expression?
- Were the works legally purchased, licensed, scraped from the open web, or taken from an unlicensed dataset?
- Does the AI product compete with the copyright owner or with a market the copyright owner could license?
- Is the source discussing a court ruling, a complaint, a settlement proposal, a government report, or an opinion piece?
Those questions are not a dodge. They are the difference between a slogan and an argument. In Ross, direct competition and unnecessary copying mattered. In the Copyright Office report, the concern sharpened around training systems that generate competing expressive content. In Anthropic, legally purchased books and unlicensed works did not receive the same treatment.[1][2][3]
This is also the kind of distinction standardized tests reward, even when the topic is not AI. GRE analytical writing asks students to examine assumptions. MCAT CARS passages often turn on the difference between what an author claims and what the evidence supports. SAT and ACT reading questions regularly punish answers that are too broad. “AI steals IP” is a real-world version of the same skill: define the claim before judging it.
So when someone says AI intellectual property theft is simple, the first move is to make it less simple in the right way. Are they talking about what went into the model, what came out of it, or both? Once that is clear, the facts start to matter: permission, copying, purpose, necessity, competition, and market effect. That is the framework students can actually use.
References
- Court Rules AI Training on Copyrighted Works Is Not Fair Use — Davis+Gilbert LLP
- Copyright Office Weighs In on AI Training and Fair Use — Skadden, May 2025
- A Tale of Three Cases: How Fair Use Is Playing Out in AI Copyright Lawsuits — Ropes & Gray, July 2025
- Copyright and Artificial Intelligence — U.S. Copyright Office
- AI and Copyright Law: What We Know — Built In
- AI Plagiarism: What Students Must Know in 2026 — Krater
Related exhibits & inventory
Verified outcomes
No verified outcomes on file for this exam yet
See Methodology for how outcome evidence is disclosed once logged.
Planners
No planner filed for this exam yet
A downloadable timeline template for this exam hasn't been published yet.
Tool verdicts
AI-tool cautions
No AI tools tested for this exam yet
No hands-on AI-accuracy logs have been filed for this exam.
Questions about this plan
Ask a question about a specific section, timeline, or citation in this plan — or flag something that needs correcting.

Comments
Join the discussion with an anonymous comment.