← All parts of this equation

Equation 9 · Part 8 · From Scripted Bots to Autonomous Agents: A History of AI Agent Architecture

Symbol s_t+1

Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)].Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big].
st+1s_{t+1}

What this part means

sts_t+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

sts_t+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action ata_t in state sts_t , temporal-difference Q-learning, updates an estimate toward a…

Read this part in the article →

Learn the underlying idea

A subscript is a label attached below a symbol. It often selects a time step, component, category, or member of a sequence.

Open the illustrated subscripts: which member of a family? guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.