← All parts of this equation

Equation 9 · Part 7 · A Chatbot Confessed to Being Built by a Company That Never Trained It

Symbol T

I(m)=1k∑q∈Q1[rm(q)∈T]I(m)=\frac{1}{k}\sum_{q\in Q}\mathbb{1}[r_m(q)\in\mathcal{T}]
T\mathcal{T}

What this part means

T is an input to the expression that computes the quantity on the left.

Its job in the formula

T is an input to the expression that computes the quantity on the left.

The passage around this formula

Sort a population of models M\mathcal{M}=\{m1m_1,…\dots,mNm_N\} into the three exposure classes just defined, MV\mathcal{M}_V , MH\mathcal{M}_H , and M∅\mathcal{M}_\varnothing . For each model m , define its trait incidence as I(m)=1k\frac{1}{k}∑q∈Q\sum_{q\in Q}1\mathbb{1}[rm(q)r_m(q)∈\inT\mathcal{T}] , where Q is the held-out probe set of size k and rm(q)r_m(q) is m ’s response to probe q . I(m) is a dimensionless rate between 0 and 1: the fraction of probes that elicit the trait. Average I(m) within each class to get IˉV\bar I_V , IˉH\bar I_H , and I0I_0=Iˉ∅\bar I_\varnothing . I0I_0 is the article’s single most important number, because it is the rate at which the trait shows up in models that never touched the source at all — the…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

The article lists its research sources here.