← Mathematical compendium

Published equation contexts

Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)]Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big]

Why this formula appears here

The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action ata_t in state sts_t , temporal-difference Q-learning, updates an estimate toward a…

Read the full article-specific guide →

Read the representative guide

QQ

Symbol Q

Q is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
sts_t

Symbol s_t

sts_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
ata_t

Symbol a_t

ata_t is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
α\alpha

Symbol α

α is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
rt+1r_{t+1}

Symbol r_t+1

rtr_t+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
γ\gamma

Symbol gamma

gamma is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
st+1s_{t+1}

Symbol s_t+1

sts_t+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)].Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big].

Equation 9 · AI Agents & Systems

From Scripted Bots to Autonomous Agents: A History of AI Agent Architecture

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action ata_t in state sts_t , temporal-difference Q-learning, updates an estimate toward a…

Equation guide → · Article →