Equation 1 · A History of AI Alignment as a Research Field
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation gives an approximation: it relates the quantities while allowing an approximation. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
≈
addition
How to interpret it
Its accuracy depends on the assumptions and range of use described in the article.
What the article says around this equation
Irving, Christiano, and Amodei’s 2018 “AI Safety via Debate” proposed one answer: train two copies of a model to argue opposing sides of a question in front of a human judge, under a length limit, with the judge deciding which debater was more helpful and honest [ 9 ] . The paper’s own framing draws an explicit analogy to computational complexity theory, arguing that a judge relying on a single unverified report is limited to adjudicating claims a bounded verifier can check directly, while a judge equipped with a structured, adversarial debate between two competing reporters can in principle adjudicate a substantially larger class of claims, since each debater is incentivized to expose flaws…
Read the full surrounding passage
Irving, Christiano, and Amodei’s 2018 “AI Safety via Debate” proposed one answer: train two copies of a model to argue opposing sides of a question in front of a human judge, under a length limit, with the judge deciding which debater was more helpful and honest [ 9 ] . The paper’s own framing draws an explicit analogy to computational complexity theory, arguing that a judge relying on a single unverified report is limited to adjudicating claims a bounded verifier can check directly, while a judge equipped with a structured, adversarial debate between two competing reporters can in principle adjudicate a substantially larger class of claims, since each debater is incentivized to expose flaws in the other’s argument that the judge alone would not have found: . That relation is the paper’s real theoretical claim, not decoration: it is the specific argument for why debate should, in principle, extend how complex a claim a time-limited human judge can reliably oversee, and the paper reports early supporting evidence from a simplified image-classification setting where a debate-trained judge’s accuracy rose from 59.4% to 88.9% relative to a judge working from limited direct evidence alone [ 9 ] .
Sources cited in the surrounding passage
These citations give research context. Read each source to check which claims it supports.
Return to A History of AI Alignment as a Research Field