A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The dashboard goes green, and a minute later it’s red again, with nothing deployed in between. Give the mechanism first and the fixes second, on both sides of the wire.
- Name the mechanism in one breath. “It came back to more load than it left, at the moment it could handle the least, and the retries that caused are now the load.” Behind that sentence: synchronized retries and a queued backlog arrive together, caches are cold, and the first instance to pass its health check takes everything and dies. The trigger is gone but load stays above capacity. The paper on metastable failures by Bronson et al. gives that state its name.
- Say what evidence would confirm it. Requests arriving at the service split by attempt number, if callers send one (gRPC sends
grpc-previous-rpc-attempts). If they don’t, compare the service’s request rate with the user-facing request rate at the edge: a ratio well above one after recovery is amplification. Add queue depth and the share of requests that finish after their caller has timed out. - Fix the callers. Capped exponential backoff with full , so clients stop moving in lockstep. A retry budget, so retries are a bounded fraction of traffic. Retries at one layer only: 3 attempts at each of 3 layers is up to 27 calls to the bottom service for one user request. A that stops calling a dependency that keeps failing.
- Fix the service, because you won’t control every caller. Shed load early and cheaply with HTTP 503 and a randomized
Retry-Afterwhen the whole service is over capacity, bound every queue, admit each caller up to its own quota and reject the excess with HTTP 429, and drop work whose caller’s deadline has already passed. - Say how you would bring it back. Cut offered load at the edge, put instances back into rotation together rather than one at a time, admit traffic in steps, and keep the autoscaler from scaling in during the outage.
The trap is stopping at “exponential backoff”. Say what backoff alone leaves broken.
To write the client side yourself, try the retry wrapper with full jitter. Backoff, jitter, retry budgets across layers and rate limiters are on the syllabus of Practical coding for FDE rounds, in Pro.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- Every client already uses exponential backoff. Why did it still happen?
- Half the callers belong to the customer and you cannot change their code. What do you do on the server?
- The service is down again right now. What do you do in the next ten minutes?
Where answers go wrong
- Answering “add exponential backoff” and stopping. Backoff without jitter keeps clients in lockstep, backoff without a budget still multiplies load, and neither helps against callers you do not control.
- Treating the second failure as bad luck or a new bug, instead of seeing that the retries themselves are now the load.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“It fell over again because it came back to more load than it left, at the moment it could handle the least. During the outage every caller retried on the same schedule, so their attempts arrived in waves. Queues upstream filled. At recovery the service faced first attempts, the retry waves and the backlog together, with cold caches and no warm connections, so each request cost more. It failed, and the failure created the next wave. Now the retries are the load, and the system stays down with the original cause fixed.”
“Put numbers on it. Normal load is 800 requests a second, and the service can take 1,000. During a two-minute outage every request makes four attempts, so offered load is about 3,200 a second. At recovery, cold caches cut capacity to maybe 600, so it’s more than 5x over, and every failure feeds the next wave. The fix has to get offered load under 600, not just spread it out. spreads the same attempts over time; it doesn’t remove any.”
“I’d confirm it by comparing the service’s request rate with user requests at the edge; a large ratio after recovery is retries. Then I’d fix it at three places.”
“Callers. Full jitter: sleep a random time between zero and the capped exponential delay, so clients that failed together retry apart (AWS Architecture Blog). Then a retry budget, because backoff still lets every request make every attempt. This is the shape gRPC’s retry throttling uses: successes earn tokens, failures spend them, and retries stop when the bucket is half empty.”
class RetryBudget:
"""Stops retries once failures pass about
token_ratio of successes, which at the
default of 0.1 is about 1 request in 11."""
def __init__(self, max_tokens: float = 100.0, token_ratio: float = 0.1):
self.max_tokens = max_tokens
self.token_ratio = token_ratio
self.tokens = max_tokens
def record_success(self) -> None:
self.tokens = min(self.max_tokens, self.tokens + self.token_ratio)
def record_failure(self) -> None:
self.tokens = max(0.0, self.tokens - 1)
def may_retry(self) -> bool:
return self.tokens > self.max_tokens / 2
“In a real client this sits behind a lock or an atomic, because every concurrent call on the channel shares it. In an outage every call fails, the bucket drains within 50 failures, and each caller falls back to first attempts only. That turns up to max_attempts times the traffic back into roughly the traffic itself. I’d retry at one layer only; the others fail fast, and a breaker per dependency fails fast once the error rate crosses a threshold, so callers stop queuing work for a service that is down.”
“The service. Some callers belong to the customer and I can’t change them, so the service has to protect itself. Admission is per caller: each API key gets its own concurrency quota, so the customer’s retry loop is rejected before it starves everyone else. Over the quota, the service rejects with HTTP 429, which costs almost nothing and tells that caller it is the one sending too much; HTTP 503 is for shedding load when the whole service is over capacity. The Retry-After it sends is randomized, say between 1 and 10 seconds, because a fixed value re-synchronizes every client that honors it. Callers propagate deadlines, and the service drops queued work whose deadline has already passed, since nobody is waiting for that answer. Then I’d hand the customer’s team the three-line client change, which is jitter, a retry cap and honoring Retry-After, with the graph of their retries during the incident, so the conversation is about data, not blame.”
“Recovery, right now. First cut offered load below capacity at the edge: a rate limit on the gateway or load balancer, and if the callers read a config flag, retries off. Then put instances back behind the load balancer together, not each in turn as it passes its health check, because the first in takes all the traffic and dies. Slow start doesn’t cover this case: on an AWS Application Load Balancer, a newly healthy target enters slow start only when at least one other healthy target is already out of it (ALB slow start mode), so after a full outage the ramp has to come from the edge limit. Raise the admission limit in steps, watching error rate and latency. Pin the autoscaler’s minimum so it doesn’t scale in while the service is down and meet the surge with too few instances.”
“The test I’d want before calling it fixed is a load test that kills the service, lets retries pile up, restores it, and checks that it stays up.”