Cortexa AI Glossary · How it learns
What is attention in AI?
From Cortexa Learn, by Cortexa Consulting. Last checked .
How a chatbot works out what "it" means, and why long documents cost more.
The trophy and the suitcase
Read this sentence: "The trophy didn't fit in the suitcase because it was too big." What was too big? The trophy. You knew straight away. Now change one word, so it ends "because it was too small," and "it" suddenly means the suitcase. You switched without even noticing. A chatbot has to make calls like that over and over, in every reply it writes. The part of an artificial intelligence (AI) language model that does this work is called attention.
One word at a time
A language model writes by predicting. It looks at everything so far and guesses what comes next, then the next piece, then the next. To guess well, it needs to know what the earlier words mean in this particular sentence. And words lean on each other. "It" means nothing until you know what it points back to. "Big" means something different for a trophy than for a house.
Weighing the words
Some words matter far more than others for working out "it." In our sentence, "trophy" and "big" matter a lot, and "the" hardly matters. Attention is how a model tells the difference. For each word, it gives the other words a score for how relevant they are, then blends in the high scorers most strongly. So the model's sense of "it" ends up carrying a lot of "trophy."1
Many views at once
A model doesn't do this just once. It runs several attention steps side by side, each free to look for something different, and it stacks them in layers so later ones build on what earlier ones found. The researchers who introduced the design looked at what those side-by-side views had learned. Many seemed to follow the grammar or meaning of a sentence, and a pair of them appeared to be matching words like "its" to the noun they referred to.2
The paper that named it
Attention was first proposed in 2014, as an add-on that helped translation software keep track of long sentences. The big change came in June 2017, when a team of researchers, most of them at Google, published a paper called "Attention Is All You Need." Their design, the transformer, dropped much of the older machinery and built itself around attention. The large language models behind today's chatbots grew out of that design.12
Why long inputs cost more
Attention has a price. Every word is weighed against every other word, so when the text gets twice as long, that part of the work grows about four times bigger. That's one reason chatbots have a limit on how much they can read at once, called a context window. It's also one reason a very long document can make a reply slower. Engineers keep finding ways to trim the cost, and the trade is still there.2
A borrowed name
The name can mislead. In people, attention means focus, even care. In a model it means a calculation: a set of weights that decides how much each word counts. Nothing is being noticed or enjoyed. A high score also doesn't tell you why a model gave the answer it did; seeing inside a model is a harder problem, with a topic of its own. Next time a chatbot handles a tricky "it" in your own writing, try changing one nearby word and see whether its reading changes too.
Works cited
- IBM, "What is an attention mechanism?" (checked )
- Vaswani et al., "Attention Is All You Need" (2017) (checked )