Equation 2 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the logistic function. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
with the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.
Sources cited in the article section
- [1] Deep Reinforcement Learning from Human Preferences ↗
- [2] Fine-Tuning Language Models from Human Preferences ↗
- [3] Learning to Summarize from Human Feedback ↗
- [4] Training Language Models to Follow Instructions with Human Feedback ↗
- [5] Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons ↗
These citations give research context. Read each source to check which claims it supports.
Return to What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness