Glossary

Cortexa AI Glossary · How it learns

How do reasoning models learn (RLVR)?

From Cortexa Learn, by Cortexa Consulting. Last checked .

Remember the answers in the back of a math workbook? Some chatbots learned to think that way.


Answers in the back

Remember a math workbook with the answers in the back? A kid could practice alone: try a problem, check the back, try again. Some of today's "thinking" artificial intelligence (AI) models were trained in much the same way, on problems whose answers a computer can check. The method has a long name, reinforcement learning with verifiable rewards (RLVR). The idea behind it is simpler than the name.

Models that work it out first

A reasoning model writes out steps before it answers, a bit like showing your work. You may have seen a chatbot pause and say it's thinking. Those steps help most on math, code, and puzzles, where one early slip spoils the answer. But laying out good steps is a skill, and the model had to learn it somewhere.

Learning by reward

One way machines learn is by reward. The model tries something, gets a score, and is nudged toward whatever scored higher. That's called reinforcement learning. You may know a version from chatbot training, where people compared answers and the model learned from their choices: reinforcement learning from human feedback (RLHF). It works. But human judging is slow and costly, and people don't always agree.1

An answer key for a judge

RLVR swaps the human judge for an automatic check, and it suits two kinds of problems especially well.

  • Math with a known answer: the final number is right or it's wrong.
  • Code with tests: the program passes them or it doesn't.

A right answer earns a reward, and a wrong one earns nothing. Nobody's opinion is involved, so the check can run again and again, at a scale no team of people could match.1

Where the name comes from

The name comes from researchers at the Allen Institute for Artificial Intelligence, who described the method in 2024 for their openly released Tülu 3 models. They took the usual reward setup and replaced the judge with a plain check. If the answer was verifiably correct, the model got a reward. If not, it got zero. It improved the models on math and on following exact instructions, the kind a computer can verify, like a word limit.13

Thinking that grew on its own

The best-known case came from the Chinese company DeepSeek. In a paper in Nature, in September 2025, its researchers described training an early version of their reasoning model with rewards alone: points for a correct final answer and for the right format, and no human examples of how to reason. As training went on, its answers grew longer. It started checking its own work and trying other approaches without being told to. At one point it began writing "wait" far more often as it reconsidered. The finished model added a small set of human-written examples, partly because the early version's writing was hard to read.2

Where there's no answer key

That training helps explain a pattern you may notice. Reasoning models are often strongest on math, code, and logic, where an answer key exists. Plenty of everyday questions have no key. Is this email too blunt? Which plan suits my family? Labs differ in how they train, and most don't publish the details, so not every reasoning model learned this way. Where no checker exists, the checking is yours. What's one answer you'd want to check by hand this week?

Works cited

  1. Lambert et al. (Allen Institute for AI), "Tülu 3: Pushing frontiers in open language model post-training" (2024) (checked )
  2. DeepSeek-AI (Guo et al.), "DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning." Nature 645 (2025) (checked )
  3. Allen Institute for AI, "Tülu 3: The next era in open post-training." (checked )