Module 08
Multi-tenancy, rate limiting & overload
- Tenant data isolation
Trace a request with a missing tenant_id filter and see which layers (gateway, service, row-level security) would have caught it.
- Noisy neighbors, quotas & fair scheduling
Push one tenant to 50× traffic, watch everyone else starve, then add quotas, weighted fair queues, and dedicated capacity.
- Head-of-line blocking
Put a slow job at the front of a shared FIFO and watch fast jobs pile up behind it until you split the queues.
- Rate limiting algorithms
Fire the same burst at token-bucket, fixed-window, sliding-log, and sliding-counter limiters side by side, comparing accepted requests, edge bursts, and memory per user.
- Backpressure, admission control & load shedding
Make producers outpace consumers, watch an unbounded queue grow, then bound it, shed low-priority work, and push back upstream.
- Circuit breakers
Tune failure thresholds and cooldown, and watch the breaker move between closed, open, and half-open while a flaky dependency recovers.
- Timeouts, deadline budgets & retry storms
Retry independently at each hop of an A → B → C chain and watch load multiply (3 × 3 × 3), then propagate a deadline budget to stop it.
- Autoscaling signals
Compare autoscaling on CPU with autoscaling on queue depth and job latency during a burst of slow, I/O-bound jobs.
- AI workload overload
Route variable-length inference jobs across scarce accelerators with batching, streaming cancellation, and per-tenant admission, tracking cost per successful request.