Path-Based Failure and Evolution Management (2004)
| Venue: | IN PROCEEDINGS OF THE INTERNATIONAL SYMPOSIUM ON NETWORKED SYSTEMS DESIGN AND IMPLEMENTATION (NSDI’04 |
| Citations: | 91 - 5 self |
BibTeX
@INPROCEEDINGS{Chen04path-basedfailure,
author = {Mike Y. Chen and Anthony Accardi and Emre Kıcıman and Jim Lloyd and Dave Patterson and Armando Fox and Eric Brewer},
title = {Path-Based Failure and Evolution Management},
booktitle = {IN PROCEEDINGS OF THE INTERNATIONAL SYMPOSIUM ON NETWORKED SYSTEMS DESIGN AND IMPLEMENTATION (NSDI’04},
year = {2004},
pages = {309--322},
publisher = {}
}
Years of Citing Articles
OpenURL
Abstract
We present a new approach to managing failures and evolution in large, complex distributed systems using runtime paths. We use the paths that requests follow as they move through the system as our core abstraction, and our "macro" approach focuses on component interactions rather than the details of the components themselves. Paths record component performance and interactions, are user- and request-centric, and occur in sufficient volume to enable statistical analysis, all in a way that is easily reusable across applications. Automated statistical analysis of multiple paths allows for the detection and diagnosis of complex failures and the assessment of evolution issues. In particular, our approach enables significantly stronger capabilities in failure detection, failure diagnosis, impact analysis, and understanding system evolution. We explore these capabilities with three real implementations, two of which service millions of requests per day. Our contributions include the approach; the maintainable, extensible, and reusable architecture; the various statistical analysis engines; and the discussion of our experience with a high-volume production service over several years.







