Glossary

Cortexa AI Glossary · Your data, rights, and the rules

Is training AI on copyrighted work allowed?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Authors and news publishers are taking artificial intelligence companies to court. The question underneath, and where it stands.


The headline

You may have seen a headline about authors or news publishers taking an artificial intelligence (AI) company to court. The details change from case to case. Underneath almost every one sits the same question: can a company build a model from books, articles, and pictures without asking the people who made them? As of October 2026, nobody has a single answer. Courts and governments are deciding it piece by piece, and country by country.

What training copies

Training is how a model learns. It processes enormous amounts of text and images to find patterns, and that usually means copying the material along the way. Much of it is gathered from the public web, and plenty of it is protected by copyright. So the legal question is whether making and using those copies needs the owner's permission.3

The test in the United States

In the United States (US), the answer turns on fair use, a rule that lets people use copyrighted work without permission in some cases. Judges weigh four factors.

  • The purpose of the use, including whether it makes something new.
  • The kind of work that was used.
  • How much of it was used.
  • The effect on the market for the original.

No one factor decides it. So the same law can give different answers in different situations.2

The Copyright Office's view

In May 2025, the US Copyright Office released a long report on AI training. It didn't give a yes or a no. It found that some training uses are likely to be fair use and some are not, depending on the facts. Harm to the market for the original weighs heavily. And where a working market for licensing the material exists, the Office said, the case for fair use gets weaker.13

The courts decide

The report is the Office's opinion. Judges make the final call, and many lawsuits are working their way through the courts. Early rulings in the US have pointed in different directions, often turning on details like how the material was collected and what the finished tool does with it. One appeal has been decided, and more are still to come. So any single ruling you read about is one step, and a higher court may see it differently.

Other countries, other rules

Other places start from different rules. In the European Union (EU), the law allows text and data mining of material that's lawfully available, but rights holders can opt out. Under the EU's AI law, companies that make large general-purpose models must respect those opt-outs and publish a summary of what they trained on. The United Kingdom held a public consultation on the question, and other countries are still weighing their own answers.45

Paying for data

Meanwhile, some money is changing hands. A number of AI companies now pay to license books, news, music, or stock photos, and some tools say they were trained only on licensed or public-domain material. The Copyright Office pointed to these deals as a sign that licensing can work, at least for some kinds of content.3

A sharper question

If you pick AI tools for a business, look at what each company says about where its training data came from. When real money or reputation is involved, a lawyer who knows copyright is the right person to ask. And the next time one of these headlines appears, try asking: which country is this in, and which part of the training is in dispute?

Works cited

  1. AI Act Service Desk (European Commission), "Article 53: Obligations for providers of general-purpose AI models." (checked )
  2. GOV.UK, "Copyright and Artificial Intelligence" (consultation) (checked )
  3. Ballard Spahr, "Third Circuit Addresses Fair Use in AI Training, But Leaves Generative AI Questions Unresolved" (October 2026) (checked )