Glossary

Cortexa AI Glossary · How it learns

Can AI learn from AI feedback (RLAIF, Constitutional AI, RLCD)?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Some chatbots are trained partly by another model checking answers against written rules.


The style guide

A newspaper keeps a style guide, and an editor checks each story against it before it runs. Some artificial intelligence (AI) training works in a similar way. The rules are written down, and another AI model does much of the checking. That can sound circular. It has been tested, though.

When people did the rating

A common way to polish a chatbot is reinforcement learning from human feedback (RLHF). People compare pairs of answers and pick the better one, and the model learns from many thousands of those choices. It works well. It's also slow and costly, and it's hard to gather enough ratings for every kind of question. So the next question was whether a model could do some of the rating instead.

Feedback from a model

Letting a model do the rating is called reinforcement learning from AI feedback (RLAIF). A model is given instructions about what makes an answer good, and it compares answers the way a person would have. Its choices then train the chatbot, just as human choices would.2

Does it work?

Researchers at Google tested that in 2023. They trained models with AI feedback and with human feedback on summaries, on helpful replies, and on harmless ones, then asked people to judge the results. For summaries and helpful replies, people rated the two about the same. For harmless replies, the AI-feedback version came out ahead. That's a handful of tasks, tested carefully. It's a good reason for interest, and a weak reason to assume it works the same everywhere.2

Constitutional AI

The best-known version is Constitutional AI, described by the AI company Anthropic in December 2022. It gave a model a written list of principles, which it called a constitution. First, the model wrote answers, critiqued them against those principles, and revised them, and it was trained on the improved versions. Then a model compared pairs of answers against the same principles, and those comparisons became the reward. The result was less harmful, with no person labeling harmful answers. And it tended to explain its objections instead of just refusing.1

Rules anyone can read

Writing the rules down has a side benefit: other people can read them. In January 2026, Anthropic published a new constitution for its Claude models in full, free for anyone to use. It says the document is used directly in training, including to create practice conversations and to rank possible responses. You don't have to agree with every rule. You can at least read them, which is harder to do with thousands of raters' private judgments.3

One more in the family

Researchers keep trying new versions. One, from 2023, is reinforcement learning from contrastive distillation (RLCD). It asks a model the same question twice: once with a prompt that encourages following the principles, and once with a prompt that encourages breaking them. The two answers come out clearly different, so they make a ready-made pair, a better one and a worse one, without anyone having to judge.4

People still decide

People haven't left the loop. They write the principles, choose which ones to use, and test the finished model to see whether it behaves the way they meant. AI feedback changes who does the repetitive rating, while the judgment about what a good answer looks like stays with people. If you could write one rule for a chatbot, what would it be?

Works cited

  1. Anthropic, "Constitutional AI: Harmlessness from AI feedback" (2022) (checked )
  2. Lee et al. (Google), "RLAIF vs. RLHF: Scaling reinforcement learning from human feedback with AI feedback" (2023) (checked )
  3. Anthropic, "Claude's new constitution" (2026) (checked )
  4. Yang et al., "RLCD: Reinforcement learning from contrastive distillation for LM alignment" (ICLR 2024) (checked )