Coordinated Checkpoint versus Message Log for Fault Tolerant MPI

15 years 12 months ago

Download www.cs.utk.edu

— Large Clusters, high availability clusters and Grid deployments often suffer from network, node or operating system faults and thus require the use of fault tolerant programming models. MPI is one of the most widely adopted programming models for high performance computing. There are several approaches for fault tolerance in an MPI environment. The automatic and transparent ones are based on either coordinated or uncoordinated checkpointing associated with a message log strategy. There are many protocols and optimizations for these approaches and several implementations have been made. However, few results of comparison between them exist. Coordinated checkpoint has the advantage of a very low overhead as long as the execution stays fault free. In contrary, uncoordinated checkpoint must be complemented by a message log protocol which adds a signiﬁcant penalty for all message transfers even for fault free executions. The drawbacks of coordinated checkpoint are the synchronization ...

Aurelien Bouteiller, Pierre Lemarinier, Gér

Real-time Traffic

CLUSTER 2003 | Cluster Computing | Fault Tolerant Programming | Message Log | Uncoordinated Checkpoint |

claim paper

» Coordinated checkpoint from message payload in pessimistic senderbased message logging

» Blocking vs nonblocking coordinated checkpointing for largescale fault tolerant MPI Protoc...

» Impact of Event Logger on Causal Message Logging Protocols for Fault Tolerant MPI

» MPICHV Project A Multiprotocol Automatic FaultTolerant MPI

» A Fault Tolerance Protocol with Fast Fault Recovery

» Scalable FaultTolerant Distributed Shared Memory

» Interconnect agnostic checkpointrestart in open MPI

Post Info
More Details (n/a)

Added	04 Jul 2010
Updated	04 Jul 2010
Type	Conference
Year	2003
Where	CLUSTER
Authors	Aurelien Bouteiller, Pierre Lemarinier, Géraud Krawezik, Franck Cappello

Comments (0)

Sciweavers

Coordinated Checkpoint versus Message Log for Fault Tolerant MPI

CLUSTER 2003 | Cluster Computing | Fault Tolerant Programming | Message Log | Uncoordinated Checkpoint |

Explore & Download

Productivity Tools

Sciweavers