Symbol b
the agent’s probability distribution over possible states before an observation.
Read this term in its guide →Published equation contexts
A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is . with a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…
the agent’s probability distribution over possible states before an observation.
Read this term in its guide →everything true about the task — every file changed, every external call made, every fact established —.
Read this term in its guide →This label says where the repeated addition, multiplication, or accumulation starts. Read its value or condition together with the article’s description of the index.
Read this term in its guide →Read it with the definitions, units, and assumptions supplied by the article.
A symbol can carry a different meaning in another article. Each occurrence keeps its own guide and term definitions.
Equation 9 · AI Agents & Systems
This equation states an equality: the expressions on both sides have the same value under the article’s assumptions.
A useful decomposition for what an agent needs to track is the belief it holds over the true state of its task, updated as new observations arrive. In the classical formalism, if b(s) is the agent’s probability distribution over possible states before an observation, a is the action just taken, o is the observation that followed, T is the environment’s transition model and O its observation model, the updated belief is . with a normalizing constant. Kaelbling, Littman and Cassandra’s original treatment of this update is also where the difficulty is made explicit: they show that the amount of memory an optimal policy needs is not bounded in advance by the size of the…