Equation 6 · What Alignment Actually Costs
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
the practical reason the labeller headcounts above are so low. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
How to interpret it
Read this expression with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
with n comparisons collected at a mean cost per comparison. The direct comparison published in the RLAIF paper puts numbers on both sides of a genuinely useful substitution. Lee and colleagues estimate that an AI-generated preference label produced with two inference passes of GPT-4 — used to correct for position bias, at an average of about 830 prompt tokens and 61 tokens of chain-of-thought rationale — costs about 0.06 US dollars per example, against about 0.67 US dollars per example for a human label purchased through a commercial annotation service pricing at roughly 0.11 US dollars per fifty words, applied to documents averaging 304 words [ 3 ] . That is better than a tenfold…
Read the full surrounding passage
with n comparisons collected at a mean cost per comparison. The direct comparison published in the RLAIF paper puts numbers on both sides of a genuinely useful substitution. Lee and colleagues estimate that an AI-generated preference label produced with two inference passes of GPT-4 — used to correct for position bias, at an average of about 830 prompt tokens and 61 tokens of chain-of-thought rationale — costs about 0.06 US dollars per example, against about 0.67 US dollars per example for a human label purchased through a commercial annotation service pricing at roughly 0.11 US dollars per fifty words, applied to documents averaging 304 words [ 3 ] . That is better than a tenfold difference in , and it is the strongest documented economic argument for the shift toward AI feedback that labs including Anthropic and Google have made in various forms. It is not, on its own, evidence that AI feedback is a free substitute: RLAIF’s own comparison measures win rate and cost together and finds the two feedback sources broadly comparable on the tasks tested, not that the human signal was redundant. What the price gap buys is a change in n that a fixed budget can afford, not a change in what a labelled comparison actually measures.
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.