Glossary

Cortexa AI Glossary · The basics

What does GPT stand for?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Three letters you've probably said out loud, and what each one tells you.


The letters

You've probably said "ChatGPT" out loud, maybe many times. The "Chat" part is plain enough. The letters on the end are less obvious. Most people never stop to ask, but they're worth a minute, because together they describe how the thing works. They stand for generative pre-trained transformer (GPT).

G is for generative

Generative means it makes something new. A lot of older artificial intelligence (AI) sorted and labeled things: spam or a real email, a cat or a dog. A generative model produces fresh text instead, one small piece at a time, each piece picked because it's likely to come next. Every word is a fresh choice, which is why the same question can get two different replies. The reply you read was written for you just then, word by word, and it wasn't pulled from a store of finished answers.1

P is for pre-trained

Pre-trained means it learned first, before anyone shaped it into a chat assistant. In that first stage the model read a huge collection of writing and practiced one task over and over: predicting what comes next. Coming first is the "pre" in the name. Its feel for language, and most of its general knowledge, come from that stage. It's also why a chatbot has a knowledge cutoff, because the reading stopped on a date and the model knows little about anything that came after it.1

T is for transformer

Transformer is the name of the design inside the model. Before it, many language systems read one word after another and could lose track of earlier words in a long passage. The transformer came from a 2017 research paper by a team at Google, and it caught on fast. Its key trick is to look at how each word in a passage relates to the others, which helps it work out what a word means in context. Most of the chatbots you've heard of are built on some version of it.12

Whose name it is

GPT is the name OpenAI gave its own family of models, starting with the first one in 2018. The name stuck. Other companies use other names. But the same ideas (making new text, learning first from a large body of writing, and the transformer design) run through most of today's chatbots. So the letters describe a common recipe as much as they name one company's product, and that helps when a new chatbot arrives with a name you've never heard before.1

Three words, most of the story

So the three words give you a short summary of how a chatbot is made.

  • Generative: it writes new text.
  • Pre-trained: it learned first from a huge amount of writing.
  • Transformer: it's built on a design that reads each word in light of the words around it.

What the letters can't tell you is what a given model was trained on, or how carefully. That's always a fair thing to ask, and a model maker's own documentation is the place to look.

Works cited

  1. IBM, "What is GPT (generative pre-trained transformer)?" (checked )
  2. Vaswani et al., "Attention Is All You Need" (2017) (checked )