A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
The loop is the easy part. Spend your time on judgment: when a retry is safe, and when it makes an outage worse.
- Ask what is being wrapped. “Before I write the loop: is this call safe to repeat?” A timeout on a request that creates something tells you nothing about whether it happened. If the operation is not idempotent and has no , the wrapper must not retry after the request was sent.
- State the policy in one line. Attempts capped, delay capped, full : sleep a uniform random time between zero and
min(max_delay, base * 2**attempt). Say why full jitter: without it, clients that failed together still retry in clusters, and jitter spreads their calls to a roughly constant rate, which cuts the total work (AWS Architecture Blog, “Exponential Backoff and Jitter”). Cap the total time too: the retry is useless once the caller has given up, so take a deadline and stop when the next sleep would pass it. - Classify errors with one rule, said aloud. “Retry what was refused before any work. Retry what might have worked only if the call is idempotent. Never retry what will fail the same way next time.” Then give the codes under each part:
- Refused before any work, so retry on any call: a connect timeout, 408 and 429. Honor
Retry-After, which can be a number of seconds or an HTTP-date (RFC 6585, section 4; RFC 9110, section 10.2.3). - Might have worked, so retry only an idempotent call or one with an idempotency key: 500, 502, 503, 504, a dropped connection and a read timeout. A gateway’s 502 or 504 says it got a bad answer or none (RFC 9110, section 15.6.3; section 15.6.5), not that the upstream did nothing. A 503 only says the server can’t handle the request right now (section 15.6.4), so one from a proxy can follow work the upstream already did.
- Will fail the same way, so never retry: 400, 401, 403, 404, 422 and a TLS certificate failure. The request or the setup is wrong, and it will be wrong next time.
- 409 means the state changed under you: re-read, then decide. Some APIs return it while a request with the same idempotency key is still in flight, and then a delayed retry is right.
- Refused before any work, so retry on any call: a connect timeout, 408 and 429. Honor
- Make time and randomness parameters. Inject
sleepand the random source, so a test runs instantly and checks every delay against its bound. - Close with the layers. If three layers each make three attempts, one user request can become
3 * 3 * 3 = 27calls to the bottom service. Retry at one layer, or give the request a retry budget.
The trap is except Exception: retry. It turns a bad request into several bad requests and hides the error the caller needed to see.
Idempotency keys, retry budgets across layers, rate limiters and testing retry logic with a fake clock are on the syllabus of Practical coding for FDE rounds, in Pro. Its first lesson, what FDE coding rounds test, is free.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- The call is a POST that creates a payment. Can you still retry it after a timeout?
- The server sends a Retry-After longer than your maximum delay. What do you do?
- This client calls a service that also retries its own dependency. What happens to load during an outage?
- How do you test the delays without waiting for them?
- This wraps a streaming LLM call. What changes?
Where answers go wrong
- Catching every exception and retrying, including validation and auth errors that will fail the same way every time.
- Backoff with no cap on delay or attempts, or jitter so small that clients still retry in lockstep.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“I’ll separate the policy from the classification, so each is testable on its own.”
import random
import time
from typing import Callable, TypeVar
T = TypeVar("T")
class RetryableError(Exception):
"""Raised by the classifier for failures worth another attempt."""
def __init__(self, msg: str, retry_after: float | None = None):
super().__init__(msg)
self.retry_after = retry_after
def retry(
fn: Callable[[], T],
*,
max_attempts: int = 5,
base: float = 0.2,
max_delay: float = 20.0,
# on the same clock as `clock`
deadline: float | None = None,
sleep: Callable[[float], None] = time.sleep,
uniform: Callable[[float, float], float] = random.uniform,
clock: Callable[[], float] = time.monotonic,
) -> T:
if max_attempts < 1:
raise ValueError("max_attempts must be at least 1")
for attempt in range(max_attempts):
try:
return fn()
except RetryableError as e:
if attempt == max_attempts - 1:
raise
jitter = uniform(0, min(max_delay, base * 2 ** attempt))
if e.retry_after is None:
delay = jitter
elif e.retry_after > max_delay:
# the server wants longer than we will wait: fail now
raise
else:
# jitter on top, or every told client returns at once
delay = e.retry_after + jitter
if deadline is not None and clock() + delay > deadline:
# the caller will have given up before the next attempt
raise
sleep(delay)
raise AssertionError("unreachable")
“Only RetryableError is caught. Everything else, including a ValueError from my own code, propagates on the first attempt. A Retry-After gets on top, or every client told to wait the same time comes back in the same instant. The deadline is the caller’s: the wrapper never sleeps past the point where nobody is waiting for the answer. The classifier decides what is retryable, and it needs to know whether the call is idempotent:”
from collections.abc import Mapping
from datetime import datetime, timezone
from email.utils import parsedate_to_datetime
import requests
# refused before any work was done
RETRY_ANY = {408, 429}
# the work may have happened
RETRY_IF_IDEMPOTENT = {500, 502, 503, 504}
class PermanentError(Exception):
"""Do not retry: the request is wrong,
or repeating it could do the work twice."""
def classify(status: int, headers: Mapping[str, str], idempotent: bool) -> None:
if status < 400:
return
if status in RETRY_ANY or (idempotent and status in RETRY_IF_IDEMPOTENT):
wait = parse_retry_after(headers.get("Retry-After"))
raise RetryableError(f"HTTP {status}", wait)
raise PermanentError(f"HTTP {status}")
def parse_retry_after(value: str | None) -> float | None:
if value is None:
return None
try:
return max(0.0, float(value))
except ValueError:
pass
try:
# the HTTP-date form
when = parsedate_to_datetime(value)
except (TypeError, ValueError):
return None
if when.tzinfo is None:
when = when.replace(tzinfo=timezone.utc)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
def call(send: Callable[[], requests.Response], idempotent: bool) -> requests.Response:
try:
resp = send()
# ConnectTimeout and SSLError subclass ConnectionError, so they go first
except requests.ConnectTimeout as e:
# never reached the server: safe on any call
raise RetryableError("connect timeout") from e
except requests.exceptions.SSLError:
# a bad certificate fails the same way every time
raise
except (requests.ConnectionError, requests.ReadTimeout) as e:
if idempotent:
raise RetryableError(type(e).__name__) from e
# a POST may have landed; only an idempotency key makes it safe
raise
classify(resp.status_code, resp.headers, idempotent)
return resp
“502 and 504 mean the gateway didn’t get a good answer, not that nothing happened, so a payment POST never retries on them. 503 goes in the same set, because a proxy can send one after the upstream already acted. If an API documents that its own 503 means the request was never processed, I’d move 503 to RETRY_ANY for that API only. 409 stays permanent in the wrapper, because the caller has to re-read before it decides; an API that returns 409 while the same is in flight gets it added to its retryable set.”
“The caller passes idempotent=True for GET, PUT and DELETE, and for a POST that carries an idempotency key. call is where transport failures go. A connect timeout never reached the server, so it’s safe for any call, but a dropped connection or a read timeout on a POST is not. The order of the except clauses matters, because in requests a ConnectTimeout and an SSLError are both ConnectionErrors. An SSLError is almost always setup, such as an expired certificate or a missing CA bundle, and that fails the same way every time, so it propagates at once. Retry-After can be seconds or a date, so the parser handles both.”
“Put together, a payment call looks like this. The key is made once, outside the retry, so every attempt carries the same key and the server can recognize the repeat:”
import uuid
session = requests.Session()
key = str(uuid.uuid4())
resp = retry(
lambda: call(
lambda: session.post(
url,
json=body,
headers={"Idempotency-Key": key},
# (connect, read) in seconds
timeout=(3.05, 10),
),
idempotent=True,
),
deadline=time.monotonic() + 30,
)
“Each attempt has its own connect and read timeouts, and the deadline stops the wrapper from starting a sleep the caller won’t wait out. That isn’t a strict end-to-end limit. In requests, the read timeout limits each wait for the next bytes from the server, not the whole response, so a slow trickle can run past the deadline. If the caller needs a hard limit, I’d shrink each attempt’s timeouts to the time left, and make the call from an async client that cancels it at the deadline.”
“The tests pin randomness and record sleeps instead of waiting:”
import pytest
def test_backoff_doubles_then_caps():
sleeps, calls = [], []
def flaky():
calls.append(1)
raise RetryableError("HTTP 503")
with pytest.raises(RetryableError):
# uniform returns its upper bound, pinning each delay to the ceiling
retry(flaky, max_attempts=6, base=0.2, max_delay=1.0,
sleep=sleeps.append, uniform=lambda lo, hi: hi)
assert len(calls) == 6
assert sleeps == pytest.approx([0.2, 0.4, 0.8, 1.0, 1.0])
def test_permanent_error_is_not_retried():
sleeps, calls = [], []
def bad():
calls.append(1)
raise PermanentError("HTTP 422")
with pytest.raises(PermanentError):
retry(bad, sleep=sleeps.append)
assert (len(calls), sleeps) == (1, [])
def test_long_retry_after_fails_at_once():
sleeps = []
def busy():
raise RetryableError("HTTP 429", retry_after=120)
with pytest.raises(RetryableError):
retry(busy, max_delay=20, sleep=sleeps.append)
assert sleeps == []
“In production I’d also log each retry with the attempt number and cause, because a wrapper that quietly succeeds on the fourth try hides a dependency that is failing, and nobody sees it until the retries run out.”
“If the wrapped call is a model API, four things change. First, a retried generation that already ran is billed again and can come back with different text, so for a model call ‘idempotent’ has to mean ‘safe to pay twice and get a different answer’, and I only pass idempotent=True when both are fine. Second, a streamed response that has already sent tokens to the user can’t be retried quietly, because the retry starts the text again. I retry only before the first token is forwarded; after that, I tell the caller the stream broke. Third, when a 429 comes from a tokens-per-minute limit, the wait it needs depends on the request’s size, so a large prompt should wait longer than a small one. I’d use the provider’s reset header if it sends one, and otherwise scale the base delay by the request’s token count. Fourth, the read timeout, which limits each gap between bytes, not the whole response. A generation that isn’t streamed sends nothing until it’s done, so the read timeout has to sit above the p99 generation time, or the wrapper cuts off good long answers and asks for them again. Streamed, the longest gap is usually the wait for the first token, so the timeout can be much shorter.”