Skip to content
Cortexa ConsultingCortexa Consulting
Cortexa Ground Truth
Explainer5 min readBy Joe Coffman
0:00 / 4:11

Tip: once it’s playing, tap any line to jump the audio there.

What “multimodal” really means for a campaign.

Multimodal AI works in more than text it sees images and hears audio. Here's where that genuinely helps a campaign, and where it's still a party trick.


“Can it just look at the moodboard?” The word for what that client is really asking about is multimodal Artificial Intelligence (AI) that works in more than text. It's the buzzword on every model launch this year, so here's the plain version: what it actually buys a campaign, and where it's still a party trick.

So what does “multimodal” actually mean?

A plain Large Language Model (LLM) the brain in a jar2 from our earlier posts could only read and write text. Multimodal means that same brain now has senses: it can look at an image, listen to audio, and increasingly speak or hand you a picture back. One model across several senses not a text model with separate tools bolted on the side.

The everyday way to picture it: we gave the capable intern eyes and ears. Same judgment, same brain but now you can show it a picture to actually look at instead of describing the picture in words, and play the client call audio instead of typing up a transcript first.

The brain in a jar could only read. Multimodal is the same brain given eyes and ears — and a voice.

What does it actually add?

Not a smarter brain more ways in and out of the same one. Think of it as the senses the model can now work in, handled together rather than stitched from four separate tools.

What “multimodal” adds

What's actually real?

  • Most of the commercial frontier models really are “omni.” OpenAI describes GPT-4o (the “o” is for “omni”) as one model that accepts any combination of text, audio, image, and video and generates text, audio, and images3, fast enough to hold a spoken back-and-forth it answers audio in about a third of a second on average. It isn't text with add-ons; it's a single model trained across the senses.
  • Some of the models were designed and built this way from day one. Google says it designed Gemini to be natively multimodal, pre-trained from the start4 across text, images, audio, video, and code rather than teaching a text-only model to see after the fact.
  • “Multimodal” isn't one fixed feature always check which senses, and which direction. Anthropic's Claude, for example, can understand and analyze images5 but is an image-understanding model that can't generate them. Reading a picture and making one are different skills; a model can have one without the other.

The “so what” for anyone serving tech clients

  • When a client asks the AI to “look at” or “listen to” something, that's a multimodal ask and it's usually the cheap, useful kind: describe this image, summarize this recording, check these ten layouts against the brand guide. Scope those first; they ship.
  • Be precise about direction. “Can it read our decks?” (understanding) is a different, easier project than “can it make our decks?” (generation) and a different model may win at each. Naming which sense, and which way, is what keeps a quote honest.

Next Reads