A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.

How to answer

The question is what your client does when it cannot tell whether the last attempt happened. Say it early: “A timeout on a create isn’t a failure. It’s an unknown.”

  1. Ask what the API gives you. An , a searchable client reference, or neither? That decides the design.
  2. Sort failures by what they tell you. A connect error never left: retry freely. A final 4xx, such as 400 or 422, means a wrong request: stop. But 408 and 429 mean wait out Retry-After and resend under the same key, and a 409 saying the key is still in flight means wait and resend under the same key, because the first request may still land. A read timeout, a reset or a 5xx after sending is ambiguous: the only hard case.
  3. Name the operation, not the attempt. Generate the key once per operation and store it with the body before the first send, so retries and restarts reuse both. Refuse the same key with a different body.
  4. With keys, resend under the same key, and say the provider’s rules aloud: key lifetime, and whether a stored error replays.
  5. Without keys, reconcile before you resend. Look the record up by your reference and create only if absent. That is safe only with one worker per operation and a lookup that sees its own writes; if the search lags, a miss proves nothing.
  6. With neither, never resend an ambiguous create. Mark it unknown and match records created since the send on amount, counterparty and time; a person confirms near-matches. Tell the customer their vendor’s API has the gap, and ask for a reference field.
  7. Give “unknown” a home. When attempts run out after an ambiguous error, or a refusal follows one, record unknown, not failed, for a reconciliation job to settle.

The trap is a new UUID inside the loop. It looks idempotent and protects nothing.

GlossaryIdempotency keyA client-supplied identifier that lets a server apply a repeated request once, making retries safe.More on Idempotency key

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • The API takes idempotency keys, and it replays a stored server error for the same key. What does your client do next?
  • Two workers pick up the same operation after a deploy. What stops both of them creating it?
  • Finance asks how many payments are in unknown right now and how old the oldest is. Where does that answer come from, and who gets paged when it grows?
  • The caller wants a yes or a no, and you have “unknown”. What do you return, and who resolves it?

Where answers go wrong

  • Generating a fresh idempotency key inside the retry loop, so the server sees every attempt as a new request.
  • Checking whether the record exists and then creating it, with nothing said about a second worker or a search index that lags.
  • Storing a refusal that follows a timed-out attempt as final, when the attempt that timed out may have landed.
  • Marking the operation failed when the retries run out, so someone reruns it later with a new key and creates the duplicate anyway.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

“First I’d ask what the API offers. You said some of our vendors take an idempotency key and some only store a free-text reference we can search by, so I’ll handle both behind one client.”