From Static Rate Limiting to Adaptive Traffic Management in Airbnb’s Key-Value Store (opens in new tab)
Airbnb evolved Mussel’s QoS system from static, per-client QPS limits into adaptive traffic management designed to maximize goodput. The newer approach accounts for the actual cost of requests, prioritizes critical workloads under stress, and detects hot keys or attack traffic before they overwhelm storage. Together, resource-aware quotas and real-time load shedding provide stronger protection against traffic spikes, uneven workloads, and DDoS-like bursts.
Why Static QPS Limits Fell Short
- Mussel is a multi-tenant key-value store serving millions of point and range reads across Airbnb.
- Its original Redis-backed limiter assigned each client a fixed requests-per-second quota.
- Requests exceeding the quota received HTTP 429 responses.
- This model worked when backend effort roughly matched request count.
- As usage grew, it could not account for:
- The difference between a cheap one-row lookup and a 100,000-row scan.
- Hot keys accessed by many clients simultaneously.
- Localized storage-shard overload that affected unrelated traffic.
- Sudden events such as bot floods, DDoS attacks, or large uploads.
Resource-Aware Rate Control
- Mussel replaced raw request counting with request units (RU), which represent estimated backend work.
- RU calculations incorporate:
- Fixed per-request overhead.
- Rows and payload bytes processed.
- Request latency, which distinguishes cached operations from disk-heavy ones.
- The system uses calibrated linear formulas for reads and writes, with weights based on compute, network, and disk-I/O measurements.
- Dispatchers debit a local token bucket according to each request’s RU cost rather than charging every request equally.
- Periodic RU refills preserve simple, static quotas while making them more proportional to actual resource consumption.
- Requests are rejected with HTTP 419 when the RU bucket is exhausted.
- Load shedding remains separate, allowing latency-based protection to react dynamically without changing the underlying quota-refill mechanism.
Load Shedding Under Sudden Stress
- RU rate limiting smooths normal traffic but may react too slowly to rapidly changing workloads.
- Mussel adds a load-shedding layer based on:
- Traffic criticality.
- A real-time latency ratio.
- A CoDel-inspired queue-management policy.
- Each dispatcher compares long-term p95 latency with short-term p95 latency.
- A ratio near 1.0 indicates stable performance; a drop toward 0.3 signals rapidly increasing latency.
- When stress crosses the threshold:
- The system raises the effective RU cost for a designated lower-priority client class.
- That class’s token bucket drains faster, causing its traffic to back off.
- If conditions worsen, the penalty expands to additional classes.
- Critical workloads, such as customer support and trust-and-safety traffic, can remain responsive while less important traffic is reduced.
- The latency estimate uses the constant-memory P² algorithm, avoiding raw sample storage and cross-node coordination.
Hot-Key Detection and DDoS Protection
- Client-level quotas cannot prevent overload when many clients request the same popular key.
- Mussel therefore detects skewed access patterns in real time.
- When duplicate requests target a hot key, the system can protect storage by:
- Serving responses from cache.
- Coalescing identical requests before they reach the backend.
- This approach protects the underlying shard whether the traffic comes from legitimate popularity, automation, or a DDoS burst.
Mussel’s experience suggests that mature multi-tenant services should move beyond fixed QPS limits. Combining resource-based accounting, priority-aware load shedding, and hot-key mitigation provides a more effective way to preserve reliability while maximizing useful work during unpredictable traffic conditions.