← Mathematical compendium

Published equation contexts

Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma \max_{a'} Q(s',a') - Q(s,a) \right]

Why this formula appears here

A model-free update never references a transition model. A canonical form, temporal-difference learning, updates a value estimate purely from sampled transitions: Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma \max_{a'} Q(s',a') - Q(s,a) \right]. Nothing in this update requires knowing or estimating p(s' ∣\mid s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model p^θ(s′∣s,a)\hat{p}_\theta(s' \mid s, a) — a “world model” — and plans against it directly:

Read the full article-specific guide →

Read the representative guide

QQ

Symbol Q

Q is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
ss

Symbol s

s is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
aa

Symbol a

a is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
α\alpha

Symbol α

α is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
rr

Symbol r

r is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →
γ\gamma

Symbol gamma

gamma is a part of this expression. Its role is fixed by the surrounding article and by the operations shown in the formula.

Read this term in its guide →

How to interpret it

Read this expression with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma \max_{a'} Q(s',a') - Q(s,a) \right]

Equation 1 · Robotics & Embodied AI

Comparing the Main Approaches to Robotics and Embodied AI

This mathematical expression combines the displayed quantities; its precise role follows from the surrounding article text.

A model-free update never references a transition model. A canonical form, temporal-difference learning, updates a value estimate purely from sampled transitions: Q(s,a)←Q(s,a)+α[r+γmax⁡a′Q(s′,a′)−Q(s,a)]Q(s,a) \leftarrow Q(s,a) + \alpha \left[ r + \gamma \max_{a'} Q(s',a') - Q(s,a) \right]. Nothing in this update requires knowing or estimating p(s' ∣\mid s, a) ; it only requires having experienced (s, a, r, s') . A model-based approach instead fits an explicit dynamics model p^θ(s′∣s,a)\hat{p}_\theta(s' \mid s, a) — a “world model” — and plans against it directly:

Equation guide → · Article →