Cortexa AI Glossary · How it learns
What is training data, and where does it come from?
From Cortexa Learn, by Cortexa Consulting. Last checked .
Authors and artists are asking whether their work trained these models. What did the models read?
The headline
Every so often there's a story about authors, artists, or news publishers asking whether their work was used to train artificial intelligence (AI). Some have gone to court. Under the legal arguments sits a plain question anyone can ask: what did these models learn from? The answer has a name. It's called training data.
What counts as data
Training data is every example a model studied while it was learning. For a chatbot, that's mostly text: web pages, books, articles, computer code, and conversations. An image generator learns from pictures paired with descriptions. A speech tool learns from recordings paired with transcripts. Each one is an example of what the model is supposed to get right.
Four kinds of source
The big developers describe a similar mix, drawn from four kinds of source.
- Public web pages, gathered by automated programs called crawlers.
- Collections the company licenses, such as news archives or stock photos.
- Material from users, where the product's settings and the law allow it.
- Data the company makes itself, including examples written by people it hires and examples produced by other models.
OpenAI and Anthropic both publish versions of this list. How much comes from each, and exactly which sites and books, is rarely spelled out.23
Why so much
Why so much? Because a model learns a pattern by seeing it many times, in many forms. A model that has seen ten ways of asking the same question handles the eleventh better. OpenAI says its datasets hold trillions of tokens, the small pieces of words a model reads. That's far more text than any person could read. More varied examples usually mean better guesses, and the quality of those examples counts too.2
What the data carries
Data carries its writers with it. If most of the text is in English, the model is usually stronger in English. If some communities wrote little online, their views and ways of speaking show up less. Old pages carry old facts. And a popular mistake, repeated across thousands of pages, can look like the truth to a model. The math is doing its job here. It learned from what people happened to write down.
The open questions
Two big questions are still being worked out. One is copyright: whether training on someone's work without permission is allowed. In May 2025 the United States Copyright Office released a report saying that some training uses will likely count as fair use, the legal exception that allows some copying, and some won't, depending on the facts. Courts are still deciding cases as of 2026. The other question is privacy: what personal details end up in training data, and what control people have over them. Both are moving, so check the date on anything you read.1
How to find out
You can look some of this up yourself. Many developers publish pages on how their models are trained. Since January 2026, a California law has required developers of generative AI offered there to post a summary of their training data, including whether it holds copyrighted material or personal information. Pick one tool you use and find its page. What does it say it learned from?45
Works cited
- United States Copyright Office, "Copyright and Artificial Intelligence, Part 3: Generative AI Training" (pre-publication version, 2025-05-09) (checked )
- OpenAI, "How ChatGPT and our foundation models are developed." (checked )
- Anthropic, "How do you use personal data in model training?" (checked )
- California Legislative Information, "AB-2013 Generative artificial intelligence: training data transparency" (chaptered 2024-09-28) (checked )
- OpenAI, "How your data is used to improve model performance." (checked )
- Anthropic, "Transparency Hub." (checked )