Reinforcement learning from human feedback (RLHF)
RLHF is a training step in which people compare or rate a model's answers, and the model is adjusted to prefer the kinds of answers people rated higher. It is a key reason chatbots are more helpful and polite than raw pre-trained models.
In one line, for a 12-year-old
RLHF is people giving an AI thumbs up or thumbs down until it learns which answers are helpful.
An example
When a chatbot shows you two answers and asks which is better, your choice may be used as feedback of this kind.
Why it matters to people
Whose preferences are used shapes the model's tone and values. It can also teach a model to sound agreeable rather than be accurate, a problem researchers call sycophancy.