← Back to article

Equation 19 · A Chatbot Confessed to Being Built by a Company That Never Trained It

What does this equation mean?

I0=Iˉ∅I_0=\bar I_\varnothing

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

Inputs and operationsbar I_varnothing
Result or conditionI_0
How to read the two sides of this formula. Follow the article passage for the meaning of each quantity.

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

I0I_0

Symbol I_0

the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.

Understand this part →

Iˉ∅\bar I_\varnothing

Symbol bar I_varnothing

bar IvI_varnothing is an input to the expression that computes the quantity on the left.

Understand this part →

=

=

The expressions on both sides represent the same quantity under the stated assumptions.

Understand this part →

See an illustrated explanation →
subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…
Read the full surrounding passage
Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the convergence baseline , the rate any theory of pure coincidence has to beat.

Read the equation in its article →

For background, read the article’s source list.

Return to A Chatbot Confessed to Being Built by a Company That Never Trained It

See this formula across 1 published context →

Browse the mathematical compendium →