← Mathematical compendium

Published equation contexts

r(x,y)  =  1 ⁣[ y=y⋆(x) ]r(x,y) \;=\; \mathbf{1}\!\left[\, y = y^{\star}(x) \,\right]

Why this formula appears here

normalizing each sampled response’s reward against the mean and standard deviation of a group of G responses to the same prompt rather than against a separately trained critic network [ 13 ] . Lambert and colleagues, building the fully open Tulu 3 post-training recipe, named the general approach explicitly, describing it as “a novel method we call Reinforcement Learning with Verifiable Rewards,” and used it alongside supervised fine-tuning and preference optimization rather than as a wholesale replacement for either [ 15 ] . The clearest large-scale demonstration is DeepSeek-R1: Guo and colleagues report training a model with reinforcement learning alone against a purely rule-based reward —…

Read the full article-specific guide →

Read the representative guide

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

r(x,y)  =  1 ⁣[ y=y⋆(x) ].r(x,y) \;=\; \mathbf{1}\!\left[\, y = y^{\star}(x) \,\right].

Equation 25 · AI Safety

The Main Technical Approaches to AI Alignment, Compared

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

normalizing each sampled response’s reward against the mean and standard deviation of a group of G responses to the same prompt rather than against a separately trained critic network [ 13 ] . Lambert and colleagues, building the fully open Tulu 3 post-training recipe, named the general approach explicitly, describing it as “a novel method we call Reinforcement Learning with Verifiable Rewards,” and used it alongside supervised fine-tuning and preference optimization rather than as a wholesale replacement for either [ 15 ] . The clearest large-scale demonstration is DeepSeek-R1: Guo and colleagues report training a model with reinforcement learning alone against a purely rule-based reward —…

Equation guide → · Article →