Metastable Failures in Distributed Systems
Distributed systems often fail spectacularly and unpredictably. They are a cause for a headache and sleepless on-call nights for way too many engineers. And this is despite lots of efforts to understand the failures, and all the tools and “best practices” we have to contain and/or prevent them. Today I want to talk about failures that may occur on a nearly healthy system with no So how does a “healthy” system fail? Naturally, something needs to happen for a failure to occur. We call...
Nathan Bronson Abutalib Aghayev Timothy Zhu Amazon Simpledb Service Disruption App Engine Incident நாதன் ப்ரோன்சன்
Source: charap.co