Learning from Reinforcement and Advice Using Composite Reward Functions

14 years 1 months ago

Download ranger.uta.edu

1 Reinforcement learning has become a widely used methodology for creating intelligent agents in a wide range of applications. However, its performance deteriorates in tasks with sparse feedback or lengthy inter-reinforcement times. This paper presents an extension that makes use of an advisory entity to provide additional feedback to the agent. The agent incorporates both the rewards provided by the environment and the advice to attain faster learning speed, and policies that are tuned towards the preferences of the advisor while still achieving the underlying task objective. The advice is converted to “tuning” or user rewards that, together with the task rewards, define a composite reward function that more accurately defines the advisor’s perception of the task. At the same time, the formation of erroneous loops due to incorrect user rewards is avoided using formal bounds on the user reward component. This approach is illustrated using a robot navigation task.

Vinay N. Papudesi, Manfred Huber

Real-time Traffic

Artificial Intelligence | Faster Learning Speed | FLAIRS 2003 | Lengthy Inter-reinforcement Times | User Reward |

claim paper

Post Info
More Details (n/a)

Added	31 Oct 2010
Updated	31 Oct 2010
Type	Conference
Year	2003
Where	FLAIRS
Authors	Vinay N. Papudesi, Manfred Huber

Comments (0)

Sciweavers

Learning from Reinforcement and Advice Using Composite Reward Functions

Artificial Intelligence | Faster Learning Speed | FLAIRS 2003 | Lengthy Inter-reinforcement Times | User Reward |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers