← Back to article

Equation 9 · A Chatbot Confessed to Being Built by a Company That Never Trained It

What does this equation mean?

I(m)=1k∑q∈Q1[rm(q)∈T]I(m)=\frac{1}{k}\sum_{q\in Q}\mathbb{1}[r_m(q)\in\mathcal{T}]

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Start with1
Divide byk
This relates toI(m)
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

II

Symbol I

a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait.

Understand this part →

mm

Symbol m

m is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Understand this part →

kk

Symbol k

k occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Understand this part →

qq

Symbol q

q appears in the bound of this sum. The bound states where the repeated operation starts, ends, or which values it includes.

Understand this part →

QQ

Symbol Q

the held-out probe set of size k.

Understand this part →

rmr_m

Symbol r_m

m ’s response to probe q.

Understand this part →

T\mathcal{T}

Symbol T

T is an input to the expression that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
fraction

fraction

Divide the expression above the line by the one below it.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

11

Numerator: 1

The complete quantity above the fraction bar.

Understand this part →

q∈Qq\in Q

Starting index or lower bound: qin Q

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Understand this part →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…
Read the full surrounding passage
Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.

Read the equation in its article →

For background, read the article’s source list.

Return to A Chatbot Confessed to Being Built by a Company That Never Trained It

See this formula across 1 published context →

Browse the mathematical compendium →