Glossary

Cortexa AI Glossary · Making pictures, voices, and video

How does AI make video?

From Cortexa Learn, by Cortexa Consulting. Last checked .

A dog surfs a perfect wave. Then the board melts. That melt is a clue.


The melting surfboard

A clip scrolls past: a dog riding a perfect wave, ears flapping, spray catching the sun. For a second you believe it. Then the surfboard melts into the water and comes back as a different board. That little glitch is a clue to how the clip was made. An artificial intelligence (AI) video model made it, and it started from noise.

Pictures in a row

A video is a stack of still pictures shown quickly, usually twenty-four or more every second. Making one good picture is hard enough. A few seconds of video needs a hundred or more, and every frame has to agree with the one before. The dog has to keep the same spots. The wave has to keep rolling the same way.

Same trick, more frames

Many video models build on the method most image generators use, called diffusion. They start from random noise and clean it up step by step, steered by your description. The difference is that they work on the whole clip together, so each frame is shaped with its neighbors in view. That's how the motion stays smooth, most of the time.1

Chunks of space and time

To learn from video, some models cut it into small pieces. OpenAI's research on its Sora model, published in 2024, described turning video into what it called spacetime patches: little blocks that cover one patch of the picture across a few moments in time. Working with those blocks lets one model learn from clips of different lengths, sizes, and shapes, much as a language model learns from small chunks of text.1

Why things melt

Keeping things steady over time is the hard part. The model doesn't track a surfboard the way you do, as one solid object that has to stay the same. It's matching patterns from frame to frame, and over a few seconds small slips can pile up. So hands change shape, writing wobbles, and a board turns into water. Longer clips give those slips more room, which is one reason many clips last only a few seconds. OpenAI's own research post lists weak spots like these, including glass that doesn't shatter the way real glass would.1

Sound joins in

Until recently, AI video was silent, and any sound was added afterward. Some newer models make the audio too. Google DeepMind's Veo 3, announced in May 2025, generates sound effects, background noise, and even spoken dialogue along with the picture. That makes a clip feel far more real. It also means the voice in a video can be made by the same model.23

Good uses, and one habit

There are good uses. The film studio Lionsgate has used video models to plan shots and sketch storyboards. In 2025 Netflix said a building collapse in its series The Eternaut was made with generative AI, about ten times faster than with the usual effects tools. But a short clip can now look real, so seeing a video doesn't settle much on its own. When a clip matters, ask where it came from. Who posted it first? Does a source you trust show the same thing? Which clip in your feed this week would you want to trace?45

Works cited

  1. OpenAI, "Video generation models as world simulators" (2024-02-15) (checked )
  2. Google DeepMind, "Veo." (checked )
  3. Google, "Fuel your creativity with new generative media models and tools" (Google I/O, 2025-05-20) (checked )
  4. Variety, "Lionsgate Takes Equity Stake in Runway AI, Plans to Draw on Existing Properties for AI-Generated Short-Form Series" (2026) (checked )
  5. Fast Company, "Netflix slipped something new into your favorite show" (July 2025) (checked )