Glossary

Cortexa AI Glossary · Trust, fakes, and safety

What is AI alignment?

From Cortexa Learn, by Cortexa Consulting. Last checked .

A chatbot's polite tone and its refusals were decided by people. How those choices get built in.


Someone decided that

Ask a chatbot for help and you'll notice it has habits. It answers in much the same warm, careful tone every time. It may politely decline one part of a request and help with the rest. None of that happened by accident. People decided how the chatbot should behave, then trained it to behave that way. The name for that work is alignment.

Matching what people meant

Artificial intelligence (AI) alignment is shaping an AI model so what it does matches what people meant it to do. Usually that means helpful, honest, and careful to avoid harm. People's values differ and change over time, so alignment involves judgment calls about which values come first. And it's never finished once and for all.1

The gap

A language model starts by learning to predict the next piece of text, from a huge amount of writing. That makes it fluent. It doesn't give it any sense of what you want from it. Ask a model fresh from that first stage for advice, and it might carry on your question with more questions, because that's what text online often does. Alignment is the work of closing the gap between predicting text and being useful.2

Saying what you mean is hard

Saying exactly what you mean is harder than it sounds. Every old story about a wish granted too literally makes the point: ask for everything you touch to turn to gold, and your lunch turns to gold too. A model given a simple goal, like "keep the user happy," could learn to flatter people instead of helping them. Topic 57 looks at that habit. So alignment work spends a lot of effort spelling out what good behavior looks like, case by case.

People's preferences

One common tool is reinforcement learning from human feedback (RLHF), from topic 27. People compare two answers and pick the better one, and the model is trained toward the kind they prefer. It's a little like learning manners by watching which replies people liked.1

Written principles

Another tool is to write the values down. Anthropic, the company that makes the Claude models, trains with a written set of principles it calls a constitution. In the method it first described in 2022, the model critiques and revises its own answers against those principles. Later, a model uses the same principles to judge which of two answers is better. Writing them down means anyone can read them and argue with them. In January 2026, Anthropic published a new, much longer version of Claude's constitution.34

Testing the result

Then people check whether it worked. Evaluations, sets of test questions with a way of scoring them, measure how a model behaves across many situations. Red teams try hard to make it misbehave. When testing turns up a gap, the training is adjusted and the tests run again. Topic 84 covers evaluations, and topic 90 covers red teams.1

Where you see it

You meet alignment choices every time you use a chatbot. The refusals, the tone, how much it hedges, and how it handles a sensitive subject were all decided by people. Then they were trained in. Different companies make different choices, so two chatbots can answer the same question in very different styles. Next time one declines something, ask it to explain why.

Works cited

  1. IBM, "What is AI alignment?" (checked )
  2. IBM Research, "What is AI alignment?" (checked )
  3. Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022-12-15) (checked )
  4. Anthropic, "Claude's new constitution" (2026-01) (checked )