← All symbols

Symbol in context

yy

This entry links every published equation using this exact notation. Its meaning may change between equations.

Meanings in context

Used in 63 equations

max⁡θ  Ex, y∼πθ[rϕ(x,y)]  −  β DKL(πθ ∥ πref).\max_\theta \; \mathbb{E}_{x,\, y\sim\pi_\theta}\big[r_\phi(x,y)\big] \;-\; \beta \, D_{\mathrm{KL}}\big(\pi_\theta \,\Vert\, \pi_{\mathrm{ref}}\big).

The Main Technical Approaches to AI Alignment, Compared · Equation 2

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Equation guide → · This term → · Article →
Ldistill=(1−α) CE(y, σ(zs))  +  α T2 CE(σ(zt/T), σ(zs/T))\mathcal{L}_{\text{distill}} = (1-\alpha)\,\mathrm{CE}\big(y,\ \sigma(z_s)\big) \;+\; \alpha\, T^2\,\mathrm{CE}\big(\sigma(z_t/T),\ \sigma(z_s/T)\big)

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared · Equation 1

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Equation guide → · This term → · Article →
y(x)=∑i∈TopK(G(x))G(x)i⋅Ei(x),Ctok≈kE⋅Ctokdense(Ntotal)y(x) = \sum_{i \in \mathrm{TopK}(G(x))} G(x)_i \cdot E_i(x), \qquad C_{\text{tok}} \approx \frac{k}{E} \cdot C_{\text{tok}}^{\text{dense}}(N_{\text{total}})

Shrink It, Train It Small, or Search for It: The Main Strategies for Small Models, Compared · Equation 29

This equation gives an approximation: it relates the quantities while allowing an approximation.

Equation guide → · This term → · Article →
LKD=α LCE(y,σ(zs))+(1−α) T2 KL(σ(zt/T) ∥ σ(zs/T))\mathcal{L}_{\mathrm{KD}} = \alpha \, \mathcal{L}_{\mathrm{CE}}\left(y, \sigma(z_s)\right) + (1-\alpha)\, T^2 \, \mathrm{KL}\left(\sigma(z_t / T) \,\Vert\, \sigma(z_s / T)\right)

How a Model Actually Gets Small Enough to Run on a Phone · Equation 7

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Equation guide → · This term → · Article →
J(θ)=Ex∼D, y∼πθ(⋅∣x)[r(x,y)]−βDKL(πθ(⋅∣x) ∥ πref(⋅∣x))J(\theta) = \mathbb{E}_{x \sim \mathcal{D},\ y \sim \pi_\theta(\cdot \mid x)}\left[r(x,y)\right] - \beta D_{\mathrm{KL}}\left(\pi_\theta(\cdot \mid x) \,\|\, \pi_{\mathrm{ref}}(\cdot \mid x)\right)

Claude, From First Principles: Training, Constitutional Methods, and What Actually Shapes a Response · Equation 6

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

Equation guide → · This term → · Article →
p(y∣x)≈∑z∈top-k(pη(⋅∣x))pη(z∣x) pθ(y∣x,z),p(y \mid x) \approx \sum_{z \in \mathrm{top}\text{-}k\left(p_\eta(\cdot \mid x)\right)} p_\eta(z \mid x) \, p_\theta(y \mid x, z),

From BM25 to Agentic Retrieval: A History of Retrieval-Augmented Generation · Equation 21

This equation gives an approximation: it relates the quantities while allowing an approximation.

Equation guide → · This term → · Article →
E[rθ(x,y)]−β E[DKL ⁣(πRL ∥ πSFT)]+γ E[log⁡πRL(x)],\mathbb{E}\big[r_\theta(x,y)\big] - \beta\,\mathbb{E}\big[D_{\mathrm{KL}}\!\left(\pi^{\mathrm{RL}} \,\|\, \pi^{\mathrm{SFT}}\right)\big] + \gamma\,\mathbb{E}\big[\log \pi^{\mathrm{RL}}(x)\big],

What RLHF Actually Optimises: Rated Agreeableness, and Where It Parts from Helpfulness · Equation 3

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

Equation guide → · This term → · Article →