← All parts of this equation

Equation 22 · Part 4 · The Main Technical Approaches to AI Alignment, Compared

=

PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw).\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}.
=

What this part means

The expressions on both sides represent the same quantity under the stated assumptions.

Its job in the formula

The equals sign connects the complete expression on the left with the complete expression on the right. Both sides must have compatible units.

The passage around this formula

They quantify how much of the gap this recovers with a metric called performance gap recovered, defined by its two boundary conditions: it equals one under perfect weak-to-strong generalization and zero when the weak-to-strong model does no better than the weak supervisor it learned from [ 12 ] . Writing fwf_w for the weak supervisor, fw→sf_{w\to s} for the strong model fine-tuned on the weak model’s labels, and fs→sf_{s\to s} for the strong model fine-tuned on ground truth as an upper-bound ceiling, those two conditions pin down PGR  =  perf(fw→s)−perf(fw)perf(fs→s)−perf(fw)\mathrm{PGR} \;=\; \frac{\mathrm{perf}(f_{w\to s}) - \mathrm{perf}(f_w)}{\mathrm{perf}(f_{s\to s}) - \mathrm{perf}(f_w)}. The paper also reports that simple additional interventions help: an auxiliary confidence loss recovered performance closer to GPT-3.5 level using only…

Read this part in the article →

Learn the underlying idea

An equals sign says that the expression on its left and the expression on its right have the same value under the stated definitions and assumptions.

Open the illustrated equality: what the equals sign claims guide →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.