← Mathematical compendium

Published equation contexts

at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g)

Why this formula appears here

Strip away the branding and every one of these systems is choosing an action from a history of observations rather than from the true state of the world, because the true state is never directly available to it. That is the founding move of decision theory under partial observability: an agent’s policy must be a function of a belief built from what it has seen, not of the world as it actually is, and the two can silently diverge [ 11 ] . Write the general shape once, for a single decision-maker at a single step: at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g). where o≤to_{\le t} is the observation history, hth_t is whatever state the system has chosen to retain, and g is the task goal. Nothing here is specific to language…

Read the full article-specific guide →

Read the representative guide

ata_t

Symbol a_t

ata_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
πθ\pi_\theta

Symbol pi_θ

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

at∼πθ(a∣o≤t, ht, g),a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g),

Equation 1 · AI Agents & Systems

Comparing the Main Approaches to AI Agent Architecture

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

Strip away the branding and every one of these systems is choosing an action from a history of observations rather than from the true state of the world, because the true state is never directly available to it. That is the founding move of decision theory under partial observability: an agent’s policy must be a function of a belief built from what it has seen, not of the world as it actually is, and the two can silently diverge [ 11 ] . Write the general shape once, for a single decision-maker at a single step: at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g). where o≤to_{\le t} is the observation history, hth_t is whatever state the system has chosen to retain, and g is the task goal. Nothing here is specific to language…

Meanings in this article

Equation guide → · Article →