Sciweavers

IPPS
2006
IEEE

Easy and reliable cluster management: the self-management experience of Fire Phoenix

14 years 6 months ago
Easy and reliable cluster management: the self-management experience of Fire Phoenix
High-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfill the new demands and challenges, cluster system software is inevitably complex. Even for experienced administrators, the management of a cluster system is an exhausting job. This paper introduces Fire Phoenix, a scalable and self-managing cluster system software that supports both scientific and commercial applications. With the self-configuring and self-healing features, much of the machine configuration and error recovery can be done automatically. Our design has been proven effective in the operations of the Dawning 4000A supercomputer, which is the biggest cluster system in China.
Zhihong Zhang, Dan Meng, Jianfeng Zhan, Lei Wang,
Added 12 Jun 2010
Updated 12 Jun 2010
Type Conference
Year 2006
Where IPPS
Authors Zhihong Zhang, Dan Meng, Jianfeng Zhan, Lei Wang, Linping Wu, Huang Wei
Comments (0)