← Back to article

Equation 1 · Comparing the Main Approaches to AI Agent Architecture

What does this equation mean?

at∼πθ(a∣o≤t, ht, g),a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g),

Read the formula alongside the article passage below. Each part has a deeper page with its role in the equation, the supporting passage and nearby citations.

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions. Read the equation part by part below; each part has a contextual explanation and a link to its mathematical background.

Read it piece by piece

ata_t

Symbol a_t

ata_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

πθ\pi_\theta

Symbol pi_θ

pi_θ is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Understand this part →

o≤to_{\le t}

Symbol o_ ≤ t

the observation history.

Understand this part →

hth_t

Symbol h_t

whatever state the system has chosen to retain.

Understand this part →

gg

Symbol g

the task goal.

Understand this part →

subscript

subscript

The lower label selects a particular version, component, or indexed member of the quantity. For example, x₀ and xₜ can be values at different positions.

Understand this part →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

What the article says around this equation

Strip away the branding and every one of these systems is choosing an action from a history of observations rather than from the true state of the world, because the true state is never directly available to it. That is the founding move of decision theory under partial observability: an agent’s policy must be a function of a belief built from what it has seen, not of the world as it actually is, and the two can silently diverge [ 11 ] . Write the general shape once, for a single decision-maker at a single step: at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g). where o≤to_{\le t} is the observation history, hth_t is whatever state the system has chosen to retain, and g is the task goal. Nothing here is specific to language…
Read the full surrounding passage
Strip away the branding and every one of these systems is choosing an action from a history of observations rather than from the true state of the world, because the true state is never directly available to it. That is the founding move of decision theory under partial observability: an agent’s policy must be a function of a belief built from what it has seen, not of the world as it actually is, and the two can silently diverge [ 11 ] . Write the general shape once, for a single decision-maker at a single step: at∼πθ(a∣o≤t, ht, g)a_t \sim \pi_\theta(a \mid o_{\le t},\ h_t,\ g). where o≤to_{\le t} is the observation history, hth_t is whatever state the system has chosen to retain, and g is the task goal. Nothing here is specific to language models; it is the generic shape of a controller. What actually distinguishes the four patterns below is not this equation but what each one does with hth_t : whether it is one growing transcript, a plan object computed once and then held fixed, a set of disjoint per-worker histories that never touch each other directly, or a position in an explicit graph whose edges may or may not be permitted to fire. The rest of this article works through each shape in turn, using the tradeoffs their own architects have written down, not a synthetic score.

Read the equation in its article →

Sources cited in the surrounding passage

These citations give research context. Read each source to check which claims it supports.

Return to Comparing the Main Approaches to AI Agent Architecture

See this formula across 1 published context →

Browse the mathematical compendium →