← All parts of this equation

Equation 1 · Part 2 · Comparing the Main Approaches to AI Agent Architecture

Symbol pi_θ

at∼πθ(a∣o≤t, ht, g),a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g),
πθ\pi_\theta

What this part means

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

Strip away the branding and every one of these systems is choosing an action from a history of observations rather than from the true state of the world, because the true state is never directly available to it. That is the founding move of decision theory under partial observability: an agent’s policy must be a function of a belief built from what it has seen, not of the world as it actually is, and the two can silently diverge [ 11 ] . Write the general shape once, for a single decision-maker at a single step: at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g). where o≤to_{\le t} is the observation history, hth_t is whatever state the system has chosen to retain, and g is the task goal. Nothing here is specific to language…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.