← All parts of this equation

Equation 22 · Part 8 · The Main Technical Approaches to AI Alignment, Compared

Numerator: perf(f_wto s) - perf(f_w)

PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw).\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}.
perf(fw→s)−perf(fw)\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)

What this part means

The complete quantity above the fraction bar.

Its job in the formula

perf(fwf_wto s) - perf(fwf_w) occurs above the fraction bar. The numerator is divided by the entire denominator below it.

The passage around this formula

They quantify how much of the gap this recovers with a metric called performance gap recovered, defined by its two boundary conditions: it equals one under perfect weak-to-strong generalization and zero when the weak-to-strong model does no better than the weak supervisor it learned from [ 12 ] . Writing fwf_w for the weak supervisor, fw→sf_{w\to s} for the strong model fine-tuned on the weak model’s labels, and fs→sf_{s\to s} for the strong model fine-tuned on ground truth as an upper-bound ceiling, those two conditions pin down PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw)\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}. The paper also reports that simple additional interventions help: an auxiliary confidence loss recovered performance closer to GPT-3.5 level using only…

Read this part in the article →

Learn the underlying idea

A fraction a/b means a divided by b. The top number is the numerator; the bottom number is the denominator, and it cannot be zero.

Open the illustrated fractions: division written vertically guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.