Sciweavers

Free Online Productivity Tools i2Speak i2Symbol i2OCR iTex2Img iWeb2Print iWeb2Shot i2Type iPdf2Split iPdf2Merge i2Bopomofo i2Arabic i2Style i2Image i2PDF iLatex2Rtf Sci2ools

170

AAAI
2004

103views Intelligent Agents» more AAAI 2004»

Stochastic Local Search for POMDP Controllers

15 years 8 months ago

Stochastic Local Search for POMDP Controllers

Download www.cs.utoronto.ca

The search for finite-state controllers for partially observable Markov decision processes (POMDPs) is often based on approaches like gradient ascent, attractive because of their relatively low computational cost. In this paper, we illustrate a basic problem with gradient-based methods applied to POMDPs, where the sequential nature of the decision problem is at issue, and propose a new stochastic local search method as an alternative. The heuristics used in our procedure mimic the sequential reasoning inherent in optimal dynamic programming (DP) approaches. We show that our algorithm consistently finds higher quality controllers than gradient ascent, and is competitive with (and, for some problems, superior to) other state-of-the-art controller and DP-based algorithms on large-scale POMDPs.

Darius Braziunas, Craig Boutilier

Real-time Traffic

AAAI 2004 | Gradient Ascent | Intelligent Agents | Low Computational Cost | Observable Markov Decision |

claim paper

Related Content

» Online Planning Algorithms for POMDPs

» Learning Without StateEstimation in Partially Observable Markovian Decision Processes

» Reinforcement Learning in POMDPs via Direct Gradient Ascent

» A POMDP approach to P300based braincomputer interfaces

» Stochastic Local Search for SMT Combining Theory Solvers with WalkSAT

» Localizing Search in Reinforcement Learning

» Bounded Finite State Controllers

» Dynamic Local Search for the Maximum Clique Problem

» Stochastic Problem Solving by Local Computation Based on SelfOrganization Paradigm

Post Info
More Details (n/a)

Added	30 Oct 2010
Updated	30 Oct 2010
Type	Conference
Year	2004
Where	AAAI
Authors	Darius Braziunas, Craig Boutilier

Comments (0)