← Mathematical compendium

Published equation contexts

ρ  =  CRLCpre\rho \;=\; \frac{C_{\mathrm{RL}}}{C_{\mathrm{pre}}}

Why this formula appears here

That fraction has moved substantially since 2022. DeepSeek’s R1 model provides the clearest recent public case, because both its pretraining and its reinforcement learning phase were disclosed in enough technical detail for an independent estimate to be built. Epoch AI’s own reconstruction puts DeepSeek-V3’s pretraining at about 5.3 million US dollars, based on 2,048 H800 GPUs run at roughly 2 US dollars per GPU-hour over the reported training schedule, and the initial R1-Zero reinforcement learning phase at about 1 million US dollars, assuming similar hardware utilisation to the pretraining run [ 11 ] . That is roughly seventeen to twenty percent of the pretraining figure — an order of…

Read the full article-specific guide →

Read the representative guide

CpreC_{\mathrm{pre}}

Symbol C_pre

CpC_pre occurs below the fraction bar. The quantity above the bar is divided by this expression; zero is excluded as a denominator.

Read this term in its guide →

How to interpret it

With a fixed numerator, increasing a nonzero denominator reduces the fraction. Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

ρ  =  CRLCpre\rho \;=\; \frac{C_{\mathrm{RL}}}{C_{\mathrm{pre}}}

Equation 11 · AI Safety

What Alignment Actually Costs

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

That fraction has moved substantially since 2022. DeepSeek’s R1 model provides the clearest recent public case, because both its pretraining and its reinforcement learning phase were disclosed in enough technical detail for an independent estimate to be built. Epoch AI’s own reconstruction puts DeepSeek-V3’s pretraining at about 5.3 million US dollars, based on 2,048 H800 GPUs run at roughly 2 US dollars per GPU-hour over the reported training schedule, and the initial R1-Zero reinforcement learning phase at about 1 million US dollars, assuming similar hardware utilisation to the pretraining run [ 11 ] . That is roughly seventeen to twenty percent of the pretraining figure — an order of…

Meanings in this article

Equation guide → · Article →