California fault lines: understanding the causes and impact of network failures

15 years 6 months ago

Download cseweb.ucsd.edu

Of the major factors affecting end-to-end service availability, network component failure is perhaps the least well understood. How often do failures occur, how long do they last, what are their causes, and how do they impact customers? Traditionally, answering questions such as these has required dedicated (and often expensive) instrumentation broadly deployed across a network. We propose an alternative approach: opportunistically mining "low-quality" data sources that are already available in modern network environments. We describe a methodology for recreating a succinct history of failure events in an IP network using a combination of structured data (router configurations and syslogs) and semi-structured data (email logs). Using this technique we analyze over five years of failure events in a large regional network consisting of over 200 routers; to our knowledge, this is the largest study of its kind. Categories and Subject Descriptors C.2.3 [Computer-Communication Net...

Daniel Turner, Kirill Levchenko, Alex C. Snoeren,

Real-time Traffic

Communications | End-to-end Service Availability | Major Factors | Modern Network Environments | SIGCOMM 2010 |

claim paper

Post Info
More Details (n/a)

Added	06 Dec 2010
Updated	06 Dec 2010
Type	Conference
Year	2010
Where	SIGCOMM
Authors	Daniel Turner, Kirill Levchenko, Alex C. Snoeren, Stefan Savage

Comments (0)

Sciweavers

California fault lines: understanding the causes and impact of network failures

Communications | End-to-end Service Availability | Major Factors | Modern Network Environments | SIGCOMM 2010 |

Explore & Download

Productivity Tools

Sciweavers