This paper introduces the RL-TOPs architecture for robot learning, a hybrid system combining teleo-reactive planning and reinforcement learning techniques. The aim of this system is to speed up learning by decomposing complex tasks into hierarchies of simple behaviours which can be learnt more easily. Behaviours learnt in this way can subsequently be re-used to solve a variety of problems, reducing the need to learn every new task from scratch. It is even possible to learn multiple behaviours simultaneously, thus making more e cient use of experience. We demonstrate these advantages in a simple simulated environment.
Malcolm R. K. Ryan, Mark D. Pendrith