Feature Construction for Reinforcement Learning in Hearts

14 years 5 months ago

Download webdocs.cs.ualberta.ca

Temporal difference (TD) learning has been used to learn strong evaluation functions in a variety of two-player games. TD-gammon illustrated how the combination of game tree search and learning methods can achieve grand-master level play in backgammon. In this work, we develop a player for the game of hearts, a 4-player game, based on stochastic linear regression and TD learning. Using a small set of basic game features we exhaustively combined features into a more expressive representation of the game state. We report initial results on learning with various combinations of features and training under self-play and against search-based players. Our simple learner was able to beat one of the best search-based hearts programs.

Nathan R. Sturtevant, Adam M. White

Real-time Traffic

CG 2006 | Computer Graphics | Game | Game Tree Search | TD Learning |

claim paper

Post Info
More Details (n/a)

Added	13 Oct 2010
Updated	13 Oct 2010
Type	Conference
Year	2006
Where	CG
Authors	Nathan R. Sturtevant, Adam M. White

Comments (0)

Sciweavers

Feature Construction for Reinforcement Learning in Hearts

CG 2006 | Computer Graphics | Game | Game Tree Search | TD Learning |

Explore & Download

Productivity Tools

Document Tools

Image Tools

Sciweavers