Glossary

Cortexa AI Glossary · How it learns

What are embeddings?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Your phone finds beach photos you never labeled. The map of numbers behind it.


"Beach"

Type "beach" into the search box of your phone's photo app, and it may find pictures you never labeled: sand, waves, a sunburned cousin. Nobody wrote "beach" on those photos. In one common way that search like this works, your photos and your word have both been turned into long lists of numbers, and the numbers for "beach" sit close to the numbers for those pictures. Those lists are called embeddings.

A list of numbers

An embedding is a list of numbers that stands for a thing: a word, a sentence, a photo, or a whole document. A trained model makes it. You give the model the thing, and it hands back the list. The lists are long. One popular family of text embedding models from OpenAI returns 1,536 or 3,072 numbers for each piece of text by default. Usually no single number means anything you could name, like "sandy" or "blue." The meaning lives in the whole list, and in how it compares with other lists.13

A map of meaning

It helps to picture the numbers as a location. Two numbers can place a point on a flat map, the way a latitude and a longitude place a town. An embedding does the same thing with hundreds or thousands of numbers, in a space nobody can draw, but the idea holds. The model is trained so that things with similar meanings end up near each other. In a model trained on both pictures and words, a photo of a dog lands near the word "dog," near other dog photos, and far from a tax form.12

Near means similar

Once everything sits on the map, "find things like this" becomes a simpler job: find what's nearby. The search words get their own list of numbers, and the system looks for the closest points. OpenAI's documentation puts it plainly. Small distances suggest that two things are closely related, and large distances suggest they aren't. So a search for "sunset at the shore" can find a photo with no caption.3

You've met one already

If you use face unlock, you've met something very like an embedding. Apple says its face unlock feature turns the camera's scan of your face into a mathematical representation, then compares it with the one it stored when you set the feature up. If the two are close enough, the phone opens. Topic 100, "How does my phone recognize my face?", walks through that one step by step.4

Where else they show up

Embeddings sit behind a lot of everyday features. A "more like this" row in a shopping or streaming app can come from finding the items closest to one you liked. And when a chatbot looks things up in a set of documents before it answers, a method called retrieval-augmented generation (RAG), embeddings are usually how it finds the passages that match your question.15

Close isn't correct

Near in meaning is a different thing from true. An embedding can tell a search that a passage is about your question. It can't tell whether the passage is right, or current, or written by someone you'd trust. So when a tool finds something "relevant," that's where checking starts. Next time your photo search surprises you, try a stranger word, like "celebration," and see what it decides belongs nearby.

Works cited

  1. IBM, "What is vector embedding?" (checked )
  2. Google for Developers, "Machine learning glossary." (checked )
  3. OpenAI, "Vector embeddings" (API documentation) (checked )
  4. Apple Support, "About Face ID advanced technology." (checked )
  5. Lewis et al., "Retrieval-augmented generation for knowledge-intensive NLP tasks" (NeurIPS 2020) (checked )