← All parts of this equation

Equation 25 · Part 1 · The Main Technical Approaches to AI Alignment, Compared

Symbol r

r(x,y)  =  1 ⁣[ y=y⋆(x) ].r(x,y) \;=\; \mathbf{1}\!\left[\, y = y^{\star}(x) \,\right].
rr

What this part means

r is part of the quantity the equation computes from the expression on the right.

Its job in the formula

r is part of the quantity the equation computes from the expression on the right.

The passage around this formula

normalizing each sampled response’s reward against the mean and standard deviation of a group of G responses to the same prompt rather than against a separately trained critic network [ 13 ] . Lambert and colleagues, building the fully open Tulu 3 post-training recipe, named the general approach explicitly, describing it as “a novel method we call Reinforcement Learning with Verifiable Rewards,” and used it alongside supervised fine-tuning and preference optimization rather than as a wholesale replacement for either [ 15 ] . The clearest large-scale demonstration is DeepSeek-R1: Guo and colleagues report training a model with reinforcement learning alone against a purely rule-based reward —…

Read this part in the article →

Learn the underlying idea

A function assigns an output to each allowed input. The expression f(x) means “apply f to x”.

Open the illustrated functions: inputs become outputs guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.