Symbol P_k
the probability that at least one clears by chance alone across k independently tested components.
Read this term in its guide →Published equation contexts
The reason this ordering matters is not procedural fussiness. Zhang and Nanda’s systematic study of activation patching found that the choice of corruption distribution and evaluation metric — decisions usually made informally, late, and sometimes after a first look at the data — can by itself change which components a patching sweep identifies as important [ 3 ] . If the metric is chosen after the sweep, on the grounds that it produced the most legible result, the investigation has stopped testing a hypothesis and started constructing one to fit the data. The same failure has a clean statistical description. Sweep enough components at a nominal per-test false-positive rate and report…
the probability that at least one clears by chance alone across k independently tested components.
Read this term in its guide →α is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →k is part of the quantity the equation computes from the expression on the right.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 3 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The reason this ordering matters is not procedural fussiness. Zhang and Nanda’s systematic study of activation patching found that the choice of corruption distribution and evaluation metric — decisions usually made informally, late, and sometimes after a first look at the data — can by itself change which components a patching sweep identifies as important [ 3 ] . If the metric is chosen after the sweep, on the grounds that it produced the most legible result, the investigation has stopped testing a hypothesis and started constructing one to fit the data. The same failure has a clean statistical description. Sweep enough components at a nominal per-test false-positive rate and report…