Human Feedback

Ask people to write the perfect answer to every possible prompt and the project quickly becomes impossible. Ask them which of two answers they prefer and they can provide useful information much faster.

Human feedback can take the form of demonstrations, labels, ratings, rankings, corrections, or critiques. Each format reveals a different slice of preference.

Have you ever been presented by two outputs from your favorite chat AI application and selected your favorite? You have contributed to a human preference dataset!

Pairwise Comparisons

Play the annotator: pick the better response

Prompt 1 of 40 of 4 rated

My smoke alarm won't stop chirping. What should I do?

This is the raw material of RLHF. Every click you just made is exactly the kind of judgment call thousands of anonymous raters make — and whatever they tend to reward is what the model learns to produce.

Feedback quality depends on who provides it, what instructions they receive, what context they can see, and how disagreement is handled. A speed-focused rater may prefer concise confidence. A domain expert may reward caveats and source quality.

Feedback also creates behavioral pressure. If raters prefer answers that agree with them, the model can learn sycophancy. If raters miss subtle errors, the model can learn to produce answers that look right during a quick review.

The Average Rater Is a Design Choice

Sampling creates the population whose preferences the model learns. Record annotator expertise and disagreement instead of compressing every conflict into one unexplained label.

Checkpoint

Why might two high-quality annotator groups train different model behavior?