← Back to article

Equation 7 · What Alignment Actually Costs

What does this equation mean?

cˉ\bar{c}

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

cˉ\bar{c}

Symbol barc

barc is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

with n comparisons collected at a mean cost cˉ\bar{c} per comparison. The direct comparison published in the RLAIF paper puts numbers on both sides of a genuinely useful substitution. Lee and colleagues estimate that an AI-generated preference label produced with two inference passes of GPT-4 — used to correct for position bias, at an average of about 830 prompt tokens and 61 tokens of chain-of-thought rationale — costs about 0.06 US dollars per example, against about 0.67 US dollars per example for a human label purchased through a commercial annotation service pricing at roughly 0.11 US dollars per fifty words, applied to documents averaging 304 words [ 3 ] . That is better than a tenfold…
Read the full surrounding passage
with n comparisons collected at a mean cost cˉ\bar{c} per comparison. The direct comparison published in the RLAIF paper puts numbers on both sides of a genuinely useful substitution. Lee and colleagues estimate that an AI-generated preference label produced with two inference passes of GPT-4 — used to correct for position bias, at an average of about 830 prompt tokens and 61 tokens of chain-of-thought rationale — costs about 0.06 US dollars per example, against about 0.67 US dollars per example for a human label purchased through a commercial annotation service pricing at roughly 0.11 US dollars per fifty words, applied to documents averaging 304 words [ 3 ] . That is better than a tenfold difference in cˉ\bar{c} , and it is the strongest documented economic argument for the shift toward AI feedback that labs including Anthropic and Google have made in various forms. It is not, on its own, evidence that AI feedback is a free substitute: RLAIF’s own comparison measures win rate and cost together and finds the two feedback sources broadly comparable on the tasks tested, not that the human signal was redundant. What the price gap buys is a change in n that a fixed budget can afford, not a change in what a labelled comparison actually measures.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to What Alignment Actually Costs

Browse the mathematical compendium →