← Mathematical compendium

Published equation contexts

PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw)\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}

Why this formula appears here

They quantify how much of the gap this recovers with a metric called performance gap recovered, defined by its two boundary conditions: it equals one under perfect weak-to-strong generalization and zero when the weak-to-strong model does no better than the weak supervisor it learned from [ 12 ] . Writing fwf_w for the weak supervisor, fw→sf_{w\to s} for the strong model fine-tuned on the weak model’s labels, and fs→sf_{s\to s} for the strong model fine-tuned on ground truth as an upper-bound ceiling, those two conditions pin down PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw)\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}. The paper also reports that simple additional interventions help: an auxiliary confidence loss recovered performance closer to GPT-3.5 level using only…

Read the full article-specific guide →

Read the representative guide

fw→sf_{w\to s}

Symbol f_wto s

fwf_wto s occurs above the fraction bar. The numerator is divided by the entire denominator below it.

Read this term in its guide →
fs→sf_{s\to s}

Symbol f_sto s

fsf_sto s occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →
perf(fw→s)−perf(fw)\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)

Numerator: perf(f_wto s) - perf(f_w)

The complete quantity above the fraction bar.

Read this term in its guide →
perf(fs→s)−perf(fw)\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)

Denominator: perf(f_sto s) - perf(f_w)

The complete quantity below the fraction bar; it must be nonzero for this division.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw).\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}.

Equation 22 · AI Safety

The Main Technical Approaches to AI Alignment, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

They quantify how much of the gap this recovers with a metric called performance gap recovered, defined by its two boundary conditions: it equals one under perfect weak-to-strong generalization and zero when the weak-to-strong model does no better than the weak supervisor it learned from [ 12 ] . Writing fwf_w for the weak supervisor, fw→sf_{w\to s} for the strong model fine-tuned on the weak model’s labels, and fs→sf_{s\to s} for the strong model fine-tuned on ground truth as an upper-bound ceiling, those two conditions pin down PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw)\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}. The paper also reports that simple additional interventions help: an auxiliary confidence loss recovered performance closer to GPT-3.5 level using only…

Meanings in this article

Equation guide → · Article →