← All parts of this equation

Equation 22 · Part 9 · Reliable AI Agents Are Control Systems, Not Chatbots

Symbol U

J(π)=Eπ[R]−λcEπ[C]−λrEπ[L]−λuEπ[U],J(\pi) = \mathbb{E}_\pi[R] - \lambda_c\mathbb{E}_\pi[C] - \lambda_r\mathbb{E}_\pi[L] - \lambda_u\mathbb{E}_\pi[U],
UU

What this part means

residual uncertainty at commitment.

Its job in the formula

U appears inside an expected value, so its contribution is averaged under the distribution or condition shown by that operator.

Where the article explains it

where R is task reward, C is computational and human-review cost, L is realized loss from harmful actions, and U is residual uncertainty at commitment.

The passage around this formula

Optimizing task success alone invites systems to spend unlimited resources or take unacceptable risks. A more useful objective is J(π)=Eπ[R]−λcEπ[C]−λrEπ[L]−λuEπ[U]J(\pi) = \mathbb{E}_\pi[R] - \lambda_c\mathbb{E}_\pi[C] - \lambda_r\mathbb{E}_\pi[L] - \lambda_u\mathbb{E}_\pi[U]. where R is task reward, C is computational and human-review cost, L is realized loss from harmful actions, and U is residual uncertainty at commitment. The coefficients are governance choices, not model parameters. Hospitals, game studios, semiconductor fabs, and personal coding projects should not assign them equally.

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

The article lists its research sources here.