← Mathematical compendium

Published equation contexts

b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s)b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s)

Why this formula appears here

A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s)b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s). with η\eta a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…

Read the full article-specific guide →

Read the representative guide

ss

Starting index or lower bound: s

This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.

Read this term in its guide →

How to interpret it

Read it with the definitions, units, and assumptions supplied by the article.

Research cited beside this formula

Published contexts (1)

A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.

b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s),b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s),

Equation 9 · AI Agents & Systems

The Hardest Unsolved Problems in AI Agent Architecture

This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.

A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is b′(s′)=η O(o∣s′,a)∑sT(s′∣s,a) b(s)b'(s') = \eta \, O(o \mid s', a) \sum_{s} T(s' \mid s, a)\, b(s). with η\eta a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…

Meanings in this article

  • bb: the agent’s probability distribution over possible states before an observation.
  • ss: everything true about the task — every file changed, every external call made, every fact established —.
  • η\eta: the normalizing constant.
  • OO: the observation model.
  • oo: the observation that followed.
  • aa: the action just taken.
  • TT: the environment’s transition model.
Equation guide → · Article →