Equation 2 · Part 1 · What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness
Symbol σ
What this part means
the logistic function.
Its job in the formula
σ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Full expression→Symbol σ→Article meaning
Where the article explains it
with the logistic function.
The passage around this formula
with the logistic function. Everything downstream inherits the assumptions in that line. Quality is one-dimensional. Preferences are transitive. Disagreement between raters is noise around a single underlying merit, not evidence of two different merits.
Learn the underlying idea
A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.
Open the illustrated variables: a letter stands for a value guide →
See this notation across published equations →
Sources cited in the article section
- [1] Deep Reinforcement Learning from Human Preferences ↗
- [2] Fine-Tuning Language Models from Human Preferences ↗
- [3] Learning to Summarize from Human Feedback ↗
- [4] Training Language Models to Follow Instructions with Human Feedback ↗
- [5] Rank Analysis of Incomplete Block Designs: The Method of Paired Comparisons ↗
These citations provide research context; check each source for the exact claim it supports.