← Mathematical compendium

Published equation contexts

at∼πθ(a∣o≤t,ht,g)a_t \sim \pi_\theta(a \mid o_{\le t}, h_t, g)

Why this formula appears here

Let the environment have a latent state sts_t , expose an observation oto_t , and accept an action ata_t . The model and its surrounding prompt produce a policy at∼πθ(a∣o≤t,ht,g)a_t \sim \pi_\theta(a \mid o_{\le t}, h_t, g). where hth_t is retained working state and g is the task objective. A tool translates the proposed action into an environmental transition; the environment returns evidence; and an evaluator determines whether the evidence supports continuing, replanning, escalating, or stopping. The policy may be a frontier model, but the agent is the entire closed loop.

Read the full article-specific guide →

Read the representative guide

πθ\pi_\theta

Symbol pi_θ

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
o≤to_{\le t}

Symbol o_ ≤ t

o_ ≤ t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

at∼πθ(a∣o≤t,ht,g),a_t \sim \pi_\theta(a \mid o_{\le t}, h_t, g),

Equation 4 · AI Agents & Systems

Reliable AI Agents Are Control Systems, Not Chatbots

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Let the environment have a latent state sts_t , expose an observation oto_t , and accept an action ata_t . The model and its surrounding prompt produce a policy at∼πθ(a∣o≤t,ht,g)a_t \sim \pi_\theta(a \mid o_{\le t}, h_t, g). where hth_t is retained working state and g is the task objective. A tool translates the proposed action into an environmental transition; the environment returns evidence; and an evaluator determines whether the evidence supports continuing, replanning, escalating, or stopping. The policy may be a frontier model, but the agent is the entire closed loop.

Meanings in this article

Equation guide → · Article →