SafetyNet: Improving the Availability of Shared Memory Multiprocessors with Global Checkpoint/Recovery

16 years 13 days ago

Download www.cs.wisc.edu

We develop an availability solution, called SafetyNet, that uses a uniﬁed, lightweight checkpoint/recovery mechanism to support multiple long-latency fault detection schemes. At an abstract level, SafetyNet logically maintains multiple, globally consistent checkpoints of the state of a shared memory multiprocessor (i.e., processors, memory, and coherence permissions), and it recovers to a pre-fault checkpoint of the system and re-executes if a fault is detected. SafetyNet efﬁciently coordinates checkpoints across the system in logical time and uses “logically atomic” coherence transactions to free checkpoints of transient coherence state. SafetyNet minimizes performance overhead by pipelining checkpoint validation with subsequent parallel execution. We illustrate SafetyNet avoiding system crashes due to either dropped coherence messages or the loss of an interconnection network switch (and its buffered messages). Using full-system simulation of a 16-way multiprocessor running ...

Daniel J. Sorin, Milo M. K. Martin, Mark D. Hill,

Real-time Traffic

Coherence Permissions | Hardware | ISCA 2002 | SafetyNet Minimizes Performance | Transient Coherence State |

claim paper

Post Info
More Details (n/a)

Added	15 Jul 2010
Updated	15 Jul 2010
Type	Conference
Year	2002
Where	ISCA
Authors	Daniel J. Sorin, Milo M. K. Martin, Mark D. Hill, David A. Wood

Comments (0)

Sciweavers

SafetyNet: Improving the Availability of Shared Memory Multiprocessors with Global Checkpoint/Recovery

Coherence Permissions | Hardware | ISCA 2002 | SafetyNet Minimizes Performance | Transient Coherence State |

Explore & Download

Productivity Tools

Sciweavers