Equation 4 · A Chatbot Confessed to Being Built by a Company That Never Trained It
What does this equation mean?
Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.
Read it piece by piece
Symbol M
M is part of the quantity the equation computes from the expression on the right.
Symbol m_1
is an input to the expression that computes the quantity on the left.
Symbol m_N
is an input to the expression that computes the quantity on the left.
=
The expressions on both sides represent the same quantity under the stated assumptions.
See an illustrated explanation →subscript
The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.
How to interpret it
Read it with the definitions, units, and assumptions supplied by the article.
What the article says around this equation
Sort a population of models =\{,,\} into the three exposure classes just defined, , , and . For each model m , define its trait incidence as I(m)=[] , where Q is the held-out probe set of size k and is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get , , and = . is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…
Read the full surrounding passage
Sort a population of models =\{,,\} into the three exposure classes just defined, , , and . For each model m , define its trait incidence as I(m)=[] , where Q is the held-out probe set of size k and is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get , , and = . is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.
For background, read the article’s source list.
Return to A Chatbot Confessed to Being Built by a Company That Never Trained It