Step 6 of 8 - Storage
Losing a storage server, and getting it back
Kill a blob server, then bring it back. The error count doesn't move either way, and when you bring it back decides whether the failure cost anything at all.
- There's no EBS side to this drill: EBS has no blob servers, and losing the disk under a node is the previous page.
t = 0 s healthy time-lapse: 1 s on screen = 2 min of cluster time
All six servers are heartbeating.
- Under-replicated objects
- 0
- Repaired
- 0 (0%)
- Copied
- 0 GB of 50 GB
- Reads served from object storage
- 0
- Errors
- 0
- Read p99
- 0.1 ms Illustrative, not measured.
| Data on one server | the estate's 300 GB of local copies across 6 servers | 50 GB |
|---|---|---|
| Objects | 50 GB at an average 106 MB per object | 483 |
| Time to suspect | blob.suspect_after: replicas stop counting; with the archive valid, nothing is replaced | 30 s |
| Time to failed | blob.failed_after: only now does up-replication start | 5 min |
| Copy rate | repair.op_rate_limit 256 MiB/s x repair.concurrency 8 x repair.max_controllers 2 | 4096 MiB/s |
| Time to repair | 50 GB at 4096 MiB/s, once the disk is failed | 12 s |
- Click a blob server to kill it.
S3 serves the reads while the replica rebuilds
S3 holds a copy, so reads fall back to it while the replica rebuilds.
How this is worked out
- This scene is a time-lapse and it's the only one on the site that is. The controller waits 5 min without a heartbeat before it will replace a replica, so a 1x drill would be five minutes of nothing. One second on screen is 2 min of cluster time, and the clock in the readout carries the real figure.
-
Every duration here comes from a rate or a threshold the cluster actually
ships with:
blob.suspect_after,blob.failed_after,repair.op_rate_limitxrepair.concurrencyxrepair.max_controllersfor the copy, andgc.zombie_min_agewithgc.batch_ratefor the cleanup. The rate is a ceiling the settings impose, so 12 s for 50 GB is the floor, not the forecast. - Bringing the server back before 5 min costs nothing at all: a suspect disk keeps its replicas while its objects still hold a valid archive copy and aren't down a second replica, because up-replicating a disk that's about to return is work the return immediately undoes. Fall below either of those and the controller stops waiting. Come back later and the repair has already happened - the returning copies are surplus, and clearing them up is background GC that no read is waiting on.