Glossary

Cortexa AI Glossary · How it answers

What is multimodal AI?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Point your phone at a menu and ask. What's happening when a chatbot reads a picture.


The menu

You're on vacation, and the menu is in a language you can't read. So you point your phone at it and ask what the third dish is. Or you take a photo of a plant in the yard and ask what it's called. The same assistant that reads your typing is now reading a picture. That's multimodal artificial intelligence (AI) at work.

What a mode is

A mode is a kind of information: text, images, sound, or video. Older AI tools usually handled one. A spam filter read text, and a photo app sorted pictures. A multimodal model can take in more than one kind, and sometimes produce more than one, in the same conversation. So you can show it a photo and ask about it in words, and it treats both as one question. That one change is why a chatbot can now help with a broken appliance or a confusing form, and also why its mistakes deserve a closer look.12

How a picture gets in

A language model works on tokens, small pieces of text. A picture has to be turned into something similar. One common way is to cut the image into a grid of small squares, called patches, and turn each patch into numbers the model can work with, much as it does with tokens. Anthropic's documentation says its Claude models see images as patches of 28 by 28 pixels. A bigger picture means more patches. That's why a very large photo is often shrunk before the model sees it, and why tiny print can get lost.23

Talking and listening

Your voice is another mode. Some voice assistants turn your speech into text, answer in text, then read the answer aloud. OpenAI said in 2024 that its earlier voice mode worked that way, as three separate models in a chain, and that the chain lost things like tone of voice and background sound. Its newer model was trained on text, images, and audio together, so it could respond to the sound itself, and much faster. Either way, the idea is the same: more than one kind of input and output, in one conversation.4

Where it slips

Seeing a picture is a long way from understanding it, and the mistakes follow a pattern. Anthropic's documentation lists several for its own models:

  • Counts can be approximate, especially with lots of small objects.
  • Small, blurry, or rotated images can be misread.
  • Where things sit in a picture is only a rough estimate.

And like any AI model, it can describe something that isn't there with complete confidence. So for anything that matters, like the dose on a medicine label or a number on a chart, read it yourself.2

A photo carries more than you meant

A photo often holds more than the thing you meant to ask about. A friend's face in the background. Your address on an envelope on the counter. A street sign out the window. Before you share a picture with any AI tool, take a second look at what else is in it, and crop out what it doesn't need. Then try it on something low-stakes, like that menu, and see what it gets right. Then ask it one thing you could check yourself.

Works cited

  1. IBM, "What is multimodal AI?" (checked )
  2. Anthropic, "Vision" (Claude documentation) (checked )
  3. Dosovitskiy et al. (Google Research), "An image is worth 16x16 words: Transformers for image recognition at scale" (2020; ICLR 2021) (checked )
  4. OpenAI, "Hello GPT-4o" (2024) (checked )