Equation 9 · A Chatbot Confessed to Being Built by a Company That Never Trained It
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol I
a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait.
Symbol m
m is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.
Symbol k
k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.
Symbol q
q appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.
Symbol T
T is an input to the expression that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
Starting index or lower bound: qin Q
This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
How to interpret it
With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Sort a population of models =\{,,\} into the three exposure classes just defined, , , and . For each model m , define its trait incidence as I(m)=[] , where Q is the held-out probe set of size k and is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get , , and = . is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…
Read the full surrounding passage
Sort a population of models =\{,,\} into the three exposure classes just defined, , , and . For each model m , define its trait incidence as I(m)=[] , where Q is the held-out probe set of size k and is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get , , and = . is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.
For background, read the article’s source list.
Return to A Chatbot Confessed to Being Built by a Company That Never Trained It