CoursesDistributed Reliability
Backups, Recovery and Chaos
A backup that has been restored, recovery to a point in time, a region failover that has been rehearsed, and the cascade a bulkhead stops before it reaches everything.
- Lessons
- 6
- Exercises
- 35
- Minutes
- 39
- 1
A Backup Is a Restore That Worked
After this lesson you can size a backup schedule from the recovery point objective, say what a backup that has never been restored is worth, and separate the copy that protects against a disaster from the one that protects against a mistake.
DebugMultiple choiceShort answer - 2
Recovery to a Point in Time
After this lesson you can recover a database to the moment before a mistake using a base backup and the write-ahead log, and decide what to do with the good writes that came after.
Multiple choiceShort answerTrace - 3
Failing Over a Region
After this lesson you can list what a region failover has to move, say why it is rehearsed on a schedule, and decide between active-passive and active-active from the RTO and the write model.
DebugCode orderMultiple choice - 4
The Cascade, and the Bulkhead
After this lesson you can read a cascading failure from one slow dependency to a whole site, and stop it with bulkheads, circuit breakers and load shedding before it spreads.
DebugMultiple choiceTrace - 5
Checkpoint: Backups, Recovery and ChaosCheckpoint
Backups, point-in-time recovery, region failover and cascades for a system you have not seen before.
DebugMultiple choiceShort answer - 6
Boss: The Region FailoverBoss
One incident carried from the first alert to the post-mortem: the question, the estimate, the decision, the fence, the TTL nobody set, the cascade in the new region, the revised runbook, and the tradeoff you have to defend.
Code orderMultiple choiceShort answerTrace