Module 10
Operational readiness
- SLOs & error budgets
Inject incidents across a month and watch the error budget burn down, with burn-rate alerts firing at the right time or too late.
- Metrics, logs & distributed tracing
Find the slow hop in a request's span waterfall, and see what you would have missed with only metrics or only logs.
- Actionable alerts, runbooks & on-call
Handle a simulated incident with noisy and with actionable alerts, following or lacking a runbook, and compare time to mitigation.
- Load & failure testing
Run a scripted load test with failure injection against any explorer's system and get a report of which invariants held.
- Canary rollout, feature flags & rollback criteria
Ramp a bad deploy from 1% to 100% and see how explicit rollback triggers limit the blast radius.
- Retention, deletion propagation & auditability
Delete a user and follow the deletion through the primary DB, replicas, caches, search indexes, backups, and event logs, finding the copies that didn't get it.