Glossary

Cortexa AI Glossary · How it learns

What is synthetic data?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Made-up data, made on purpose. Why it helps, and why it still gets checked.


The drill

A fire drill is a made-up emergency. Nobody's building is burning, but everyone practices as if it were, so they know what to do on the day it counts. Pilots do something similar in flight simulators, practicing failures they hope never to meet in the air. Synthetic data works on the same idea. It's data made on purpose by computers, and it's used to train or test an artificial intelligence (AI) model when the real thing is hard to get.

What it is

Most training data is collected from the world: real photos, real writing, real measurements. Synthetic data is generated by software instead. International Business Machines (IBM) describes it as information created by computer simulations or algorithms that copies some of the structure and statistics of real data. It can be pictures, video, text, or rows in a table. Because a computer makes it, a team can make as much as it needs, in exactly the situations it needs. And to a model in training, a good synthetic example can look a lot like a real one.1

Why people make it

People reach for synthetic data when the real kind is hard to come by.

  • Some moments are rare or risky to capture, like a deer stepping onto a dark road in front of a self-driving car.
  • Some data is private, like customer records, and made-up examples with the same patterns can stand in for it.
  • Some data is slow and costly to label by hand, while made-up data can arrive with its labels already attached.

IBM Research names all three, and adds a plainer reason: it's cheap to produce.12

How it's made

There are a few ways to make it. A simulation can build a virtual world, the way a video game engine renders endless streets in rain and fog. Simple rules can churn out realistic but fake names and addresses for testing software. And more and more often, one AI model writes examples for another: practice questions, sample conversations, worked math problems.12

The catch

Synthetic data can only be as good as whatever made it. If the simulation never puts a cyclist on a dark street, the data won't have one, and a model trained only on that data may never have seen one. If the model writing the examples has a blind spot or a bias, its examples carry that along. IBM points out that synthetic data can inherit the biases of the real data behind it. Data can also look realistic without being accurate, and a model in training has no way to tell.3

The care it needs

So careful teams check it like any other source. They review samples by hand, mix synthetic data with real data instead of swapping one for the other, and test the finished model on real cases before trusting it. Then they keep checking. IBM notes that keeping real data in the mix helps avoid a known problem, where models trained again and again on AI-made data get worse. When a company says its model learned from synthetic data, you can ask how they tested it on the real thing.3

Works cited

  1. IBM, "What is synthetic data?" (checked )
  2. IBM Research, "What is synthetic data?" (checked )
  3. IBM, "Examining synthetic data: The promise, risks and realities." (checked )