Skip to content
Live
Loading the latest AI news…

Glossary · How models work

Reinforcement learning from human feedback (RLHF)

Also called RLHF

RLHF is a training step in which people compare or rate a model's answers, and the model is adjusted to prefer the kinds of answers people rated higher. It is a key reason chatbots are more helpful and polite than raw pre-trained models.

In one line, for a 12-year-old

RLHF is people giving an AI thumbs up or thumbs down until it learns which answers are helpful.

An example

When a chatbot shows you two answers and asks which is better, your choice may be used as feedback of this kind.

Why it matters to people

Whose preferences are used shapes the model's tone and values. It can also teach a model to sound agreeable rather than be accurate, a problem researchers call sycophancy.