← All parts of this equation

Equation 9 · Part 3 · From Scripted Bots to Autonomous Agents: A History of AI Agent Architecture

Symbol a_t

Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)].Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big].
ata_t

What this part means

ata_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

ata_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

…each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action ata_t in state sts_t , temporal-difference Q-learning, updates an estimate toward a bootstrapped target rather than waiting for a final outcome: Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)]Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big]. That equation is the real content of the RL-era architecture, not decoration: the agent’s decision rule is now a function fit…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.