Continuous-Time Hierarchical Reinforcement Learning

15 years 10 days ago

Download www.cs.ualberta.ca

Hierarchical reinforcement learning (RL) is a general framework which studies how to exploit the structure of actions and tasks to accelerate policy learning in large domains. Prior work in hierarchical RL, such as the MAXQ method, has been limited to the discrete-time discounted reward semiMarkov decision process (SMDP) model. This paper generalizes the MAXQ method to continuous-time discounted and average reward SMDP models. We describe two hierarchical reinforcement learning algorithms: continuous-time discounted reward MAXQ and continuous-time average reward MAXQ. We apply these algorithms to a complex multiagent AGV scheduling problem, and compare their performance and speed with each other, as well as several well-known AGV scheduling heuristics.

Mohammad Ghavamzadeh, Sridhar Mahadevan

Real-time Traffic

Average Reward Maxq | Discrete-time Discounted Reward | ICML 2001 | Machine Learning | MAXQ Method |

claim paper

Post Info
More Details (n/a)

Added	17 Nov 2009
Updated	17 Nov 2009
Type	Conference
Year	2001
Where	ICML
Authors	Mohammad Ghavamzadeh, Sridhar Mahadevan

Comments (0)

Sciweavers

Continuous-Time Hierarchical Reinforcement Learning

Average Reward Maxq | Discrete-time Discounted Reward | ICML 2001 | Machine Learning | MAXQ Method |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers