Symbol Q
Q is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Published equation contexts
The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action in state , temporal-difference Q-learning, updates an estimate toward a…
Q is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →α is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →gamma is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →+1 is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.
Read this term in its guide →Read this expression with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 9 · AI Agents & Systems
This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.
The next architectural break replaced the hand-written operator library with something the agent acquired through trial, error, and a numeric reward signal, rather than something a person specified in advance. Sutton and Barto’s textbook formalizes the setting that every reinforcement-learning agent since has used: an agent and an environment exchanging a state, an action, and a scalar reward at each discrete time step, with the agent’s goal defined as maximizing cumulative reward rather than satisfying a hand-specified goal predicate [ 3 ] . The canonical learning rule for estimating the value of taking action in state , temporal-difference Q-learning, updates an estimate toward a…
Equation guide → · Article →