Glossary

Cortexa AI Glossary · How it learns

What is RLHF?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Two answers side by side, and a question: which do you prefer? That choice trains chatbots.


Two answers, pick one

Now and then a chatbot shows you two answers side by side and asks which one you prefer. It's a small screen, easy to skip, and you may have tapped one without a second thought. But that kind of choice, made over and over by many people, is the raw material of a training method with a long name.

The name, word by word

It's called reinforcement learning from human feedback (RLHF), and each half of the name means something plain.

  • Reinforcement learning is training by trial and reward: a system tries something, gets a score, and does more of what scored well, the way a dog learns a trick from treats.
  • Human feedback means the scores start with people's judgments about which answers are better.

Put together, RLHF teaches a chatbot to give the kind of answers people tend to prefer.1

People rank

It starts with people. A model writes several answers to the same request, and trained raters compare them, judging which is more helpful, more accurate, or less harmful. Comparing is easier than scoring. Most of us can't say whether an answer deserves a six or a seven out of ten, but we can usually say which of two is better. Tens of thousands of these judgments become the training material.13

A scorer learns their taste

People can't rate every answer a chatbot will ever write, so the next step builds a stand-in. Engineers train a second model, called a reward model, on all those comparisons. Its job is to look at an answer and predict how much people would like it, a bit like a critic who has studied every review a restaurant ever got. It becomes a judge that never sleeps, reflecting the raters' taste as well as it can, and no better.13

The chatbot chases the score

Then the main model practices. It writes answers, the reward model scores them, and the chatbot's numbers are nudged toward whatever earns higher scores. Over many rounds it learns to produce the kind of answer people tended to pick. Engineers also hold it close to its earlier self, so it doesn't drift into strange answers that happen to fool the scorer.3

What it bought

It worked. In OpenAI's 2022 InstructGPT study, people preferred answers from a model tuned this way over answers from an untuned one more than a hundred times its size. The tuned models also made up facts less often in OpenAI's tests. A good share of the helpful, polite, on-topic manner people expect from chatbots traces back to this step.23

What it can cost

There's a catch. People often like answers that agree with them, and a scorer trained on their choices can pick up that taste too. Researchers at Anthropic reported in 2023 that people, and the reward models trained on their choices, sometimes preferred a convincing, agreeable answer over a correct one. So a chatbot tuned this way can lean toward telling you what you want to hear. Next time one agrees with you straight away, try asking it what you might be getting wrong.45

Works cited

  1. IBM, "What is reinforcement learning from human feedback (RLHF)?" (checked )
  2. OpenAI, "Aligning language models to follow instructions" (2022) (checked )
  3. Ouyang et al., "Training language models to follow instructions with human feedback" (2022) (checked )
  4. Sharma et al. (Anthropic), "Towards understanding sycophancy in language models" (2023) (checked )
  5. OpenAI, "Expanding on what we missed with sycophancy" (2025) (checked )