← Back to article

Equation 15 · A Chatbot Confessed to Being Built by a Company That Never Trained It

What does this equation mean?

I(m)I(m)

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

II

Symbol I

a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait.

Understand this part →

mm

Symbol m

m is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…
Read the full surrounding passage
Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.

Read the equation in its article →

For background, read the article’s source list.

Return to A Chatbot Confessed to Being Built by a Company That Never Trained It

See this formula across 4 published contexts →

Browse the mathematical compendium →