← Mathematical compendium

Published equation contexts

at∼πθ(a∣o≤t,ht,g),ot+1=E(st,at)a_t\sim\pi_\theta(a\mid o_{\le t},h_t,g), \qquad o_{t+1}=E(s_t,a_t)

Why this formula appears here

An agent trajectory can be represented as at∼πθ(a∣o≤t,ht,g),ot+1=E(st,at)a_t\sim\pi_\theta(a\mid o_{\le t},h_t,g), \qquad o_{t+1}=E(s_t,a_t). where the model policy chooses an action, environment E executes it against state sts_t , and the result becomes the next observation. This closes a feedback loop absent from one-shot completion.

Read the full article-specific guide →

Read the representative guide

o≤to_{\le t}

Symbol o_ ≤ t

o_ ≤ t is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →
hth_t

Symbol h_t

hth_t is an argument of the function-like quantity on the left; its role is set by that function’s stated inputs.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

at∼πθ(a∣o≤t,ht,g),ot+1=E(st,at),a_t\sim\pi_\theta(a\mid o_{\le t},h_t,g), \qquad o_{t+1}=E(s_t,a_t),

Equation 6 · History of Technology

From Autocomplete to Delegation: The Technical History Behind OpenAI Codex

This equation states a bound: one expression must stay on the indicated side of the other under the article’s assumptions.

An agent trajectory can be represented as at∼πθ(a∣o≤t,ht,g),ot+1=E(st,at)a_t\sim\pi_\theta(a\mid o_{\le t},h_t,g), \qquad o_{t+1}=E(s_t,a_t). where the model policy chooses an action, environment E executes it against state sts_t , and the result becomes the next observation. This closes a feedback loop absent from one-shot completion.

Meanings in this article

  • sts_t: the environment e executes it against state.
Equation guide → · Article →