Symbol c
c is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →Published equation contexts
The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: . This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…
c is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Read this term in its guide →m is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →M is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →orrupt is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →lean is one of the signed contributions combined to compute the quantity on the left.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 20 · AI Research
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
The standard instrument is activation patching. Run the model once on a clean input and once on a deliberately corrupted variant of it — a single token changed, a name swapped — caching every intermediate activation from both runs. Then run the model a third time on the corrupted input, but with one chosen component’s activation overwritten by the value it took during the clean run, and measure the resulting change in some scalar behavioural metric m , typically the gap between the correct answer’s logit and a specific incorrect competitor’s logit: . This particular direction — a clean value spliced into an otherwise corrupted run — is called denoising: it measures how…
Equation guide → · Article →