Distributed Reliability

Keep a system correct when parts of it fail: replication and the consistency it can promise, coordination and leases, transactions across services, and the recovery that gets the data back.
- Islands
- 4
- Lessons
- 24
- Exercises
- 140
- Island 1
Replication Models
One leader, several leaders or none; the quorum arithmetic that makes a leaderless read see the last write; the conflict two leaders create and the ways to resolve it; and the guarantees a session can be given.
- Island 2
Coordination and Leases
Why a lock across machines is a lease, the fencing token that makes a stale holder harmless, how a leader is elected and what it may assume, and how far a clock can be trusted.
- Island 3
Transactions Across Services
A saga of local transactions with compensations, the state machine that lets it resume after a crash, why two-phase commit is the last resort, and the anomalies a saga lets through that a transaction would not.
- Island 4
Backups, Recovery and Chaos
A backup that has been restored, recovery to a point in time, a region failover that has been rehearsed, and the cascade a bulkhead stops before it reaches everything.