← All parts of this equation

Equation 9 · Part 7 · From Scripted Bots to Autonomous Agents: A History of AI Agent Architecture

Symbol a

Q(st,at)←Q(st,at)+α[rt+1+γmax⁡aQ(st+1,a)−Q(st,at)].Q(s_t,a_t) \leftarrow Q(s_t,a_t) + \alpha\Big[r_{t+1} + \gamma \max_{a} Q(s_{t+1},a) - Q(s_t,a_t)\Big].
aa

What this part means

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Its job in the formula

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

The passage around this formula

The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with…

Read this part in the article →

Learn the underlying idea

A variable is a named place for a value. Its letter is a local label: x can mean position in one formula and a data point in another.

Open the illustrated variables: a letter stands for a value guide →

See this notation across published equations →

Sources cited in the surrounding passage

These citations provide research context; check each source for the exact claim it supports.