In this post12 sections
  1. Why FDE interviews use flaky APIs
  2. Prompts employers have published and candidates report
  3. Questions to ask about the API before you code
  4. Retries: exponential backoff with full jitter
  5. Which errors not to retry
  6. Cursor pagination that survives a restart
  7. Idempotency keys and the retried POST
  8. How to narrate it so the interviewer hears judgment
  9. Practice it before it counts
  10. Questions people ask
  11. Keep reading
  12. More from the blog

You open the prompt and there it is: an API you did not write, which pages its results, throttles you after a burst of calls and throws an HTTP 500 when it feels like it, with a timer running in the corner. That is the ’s job in miniature: a system you do not control, a customer who needs the data, and a deadline. An API integration coding interview checks whether you retry only what is safe to retry, with exponential backoff and ; page with a cursor you save, so a crash does not start you over; and make every write idempotent, so a retry never creates a duplicate. This post works each habit in short, runnable Python, with the words to say while you type. For the whole loop, round by round, start with the FDE interview guide.

Why FDE interviews use flaky APIs

An FDE spends the job inside somebody else’s systems: the customer’s permit database, their CRM, their approval spreadsheet, their vendor’s half-documented REST endpoint. A flaky API task shows the difference between two kinds of code:

  • Happy-path code calls the endpoint in a loop, appends results to a list and prints them. It works in the demo and falls over on the first HTTP 503.
  • Integration code assumes the network will fail, knows which failures are worth another try, remembers where it got to, and never writes the same record twice.

Our view, not a published rubric: the round is less about clever algorithms and more about whether you expect failure before it happens. The lesson on what coding rounds test covers the wider coding round; this post is the integration slice of it.

Prompts employers have published and candidates report

Some FDE take-homes are public on GitHub, and some candidates have described API tasks.

Published by the employer:

  • A deliberately broken permit API. A repo named fde-take-home, created in May 2025 in the Forerunner Industries GitHub organization, holds a mock third-party permit-system API built with intentional issues, including rate limiting after 5 consecutive requests within 10 seconds and a 20% chance of an HTTP 500 on any request. Source 1Permit System APIPublisherForerunner Industries (GitHub organization)Source typecompany website Its README gives no task text, so it does not say what candidates are asked to build.
  • A two-way spreadsheet sync. Coframe’s public forward deploy take-home gives candidates 2 hours to build a two-way integration between a client’s Google Sheet approval tracker and a mock Coframe variant-approval API, and asks for code plus a recorded video covering a demo, the solution (including every AI tool used), edge cases and productionization. Source 2Coframe Forward Deploy — Take-HomePublisherCoframe (GitHub)Source typecompany website
  • Create, then poll. Notch published a Forward Deployed Engineer home assignment in which one TypeScript task is to create a return label against a create-then-poll-status API, and its README says: “Think about how you’d handle waiting, retries, failures, and timeouts.” Source 3fde-task: Forward Deployed Engineer Home AssignmentPublisherNotch (get-notch on GitHub)Source typecompany website
  • A contract with retries written in. A vercel-solutions repo, fde-challenge-backend, specifies a service that forwards valid events to a partner catalog API with the event id as the , treats HTTP 429 and HTTP 503 as transient with at most three attempts in total, and keeps at most four catalog calls in flight. Source 4HTTP contractPublisherVercel (vercel-solutions on GitHub)Source typecompany website The file does not itself say it is an interview.

Reported by candidates:

None of this says how often an API task appears. It does show the vocabulary: rate limits, random server errors, transient status codes, idempotency keys, retries and timeouts.

Questions to ask about the API before you code

Say them out loud, and write the answers as a comment at the top of your file.

Ask about the contract first

  • How does it page: a cursor, an offset or page numbers? Can records be added while I page?
  • How does it tell me I am going too fast: HTTP 429, a Retry-After header, quota headers?
  • Which calls write data, and does the API accept an idempotency key on them?
  • What does failure look like: status codes, timeouts, or an HTTP 200 with an error in the body?
  • Does the auth token expire during a long run?
  • When I give up on a record, what should happen: fail the job, skip it and report it, or alert someone?
  • Is this a one-off or a sync that runs forever? For a two-way sync, which side wins a conflict?

The words to use: “Before I write anything, I want to pin down the contract. How does it page, how does it rate-limit, and can I send an idempotency key on writes?”

If the interviewer says “assume whatever you like”, don’t freeze. Pick, and say it: “Then I’ll assume cursor pagination, HTTP 429 with a Retry-After header, and an Idempotency-Key header on POST. I’ll keep each one behind a small function so it’s cheap to change.”

Retries: exponential backoff with full jitter

Here is a retry wrapper, standard library only. send makes one request and returns a status, headers and a body. It also translates the library’s errors: with urllib, return an HTTPError’s code and headers and raise ConnectionError for a URLError; with requests, map requests.Timeout and requests.ConnectionError the same way. In the room, say: “I’m assuming send normalizes transport errors.” The wrapper decides whether to try again and how long to wait.

import random
import time

RETRY_STATUS = {408, 429,
                500, 502, 503, 504}

def backoff(attempt, base=0.5, cap=20.0):
    # full jitter: 0 to the capped value
    return random.uniform(
        0, min(cap, base * 2 ** attempt))

def with_retries(send, attempts=5,
                 cap=20.0,
                 sleep=time.sleep):
    for attempt in range(attempts):
        try:
            status, headers, body = send()
        except (TimeoutError,
                ConnectionError):
            status, headers, body = (
                None, {}, None)
        ok = status not in RETRY_STATUS
        if status is not None and ok:
            # not retryable: return it
            return status, body
        if attempt == attempts - 1:
            raise RuntimeError("gave up")
        wait = headers.get("Retry-After")
        if wait:
            sleep(min(float(wait), cap))
        else:
            sleep(backoff(attempt, cap=cap))

Four choices to say out loud:

  1. The ceiling doubles, up to a cap. That is exponential backoff: a struggling server gets room to recover, and the cap keeps any single wait bounded.
  2. The wait is random. With full jitter it is anywhere from zero to the capped value. Without jitter, clients that failed together retry together, and the recovering service takes synchronized waves. The AWS Architecture Blog post Exponential Backoff And Jitter compares the variants if an interviewer asks why full jitter.
  3. The server’s hint wins. When a response carries Retry-After, the wrapper waits that long instead. RFC 9110, section 10.2.3 defines the header, and it can be a date as well as a number of seconds; parse the date form in production, and cap the wait either way, as the code does, so a Retry-After: 3600 can’t stall the job for an hour.
  4. sleep is a parameter. Tests pass a function that records the wait instead of sleeping, so the retry path runs in milliseconds and you can see what it did.

The test, with a stub that fails twice and then succeeds:

replies = iter([
    (503, {}, None),
    (429, {"Retry-After": "2"}, None),
    (200, {}, {"ok": True}),
])
waits = []
print(with_retries(lambda: next(replies),
                   sleep=waits.append))
# (200, {'ok': True})
print(waits)
# [0.32..., 2.0]
# first wait random, 0 to 0.5;
# second from the header

A stub that returns HTTP 400 gets one call and its error back, and one that always raises ConnectionError ends in RuntimeError. Run both to show the wrapper stops.

To practice saying this aloud, try the question on why backoff needs jitter. It is free.

Which errors not to retry

A retry is a bet that the same request will get a different answer. Place it only when the failure is about timing, not the request.

ResponseWhat to do
Timeout, dropped connection or HTTP 408Retry a read. Retry a write only with an idempotency key
HTTP 429Retry after Retry-After, or back off
HTTP 500, HTTP 502, HTTP 503 or HTTP 504Retry a read with backoff. Retry a write only with an idempotency key
HTTP 400 or HTTP 422Don’t retry. Log the record and move on
HTTP 401Refresh the token once, then retry once
HTTP 403 or HTTP 404Don’t retry. Report it
HTTP 409Don’t retry blindly. Read the current state first

Three notes to go with the table:

  • HTTP 429 is a request to slow down, defined in RFC 6585, section 4. If you keep hitting it, your client is too fast, and retries alone won’t fix that. Pace the calls; the token bucket rate limiter post builds the limiter for that.
  • A timeout is the dangerous one. You don’t know whether the server did the work. For GET that doesn’t matter: RFC 9110, section 9.2.2 defines GET, PUT and DELETE as idempotent, so repeating them leaves the server in the same state. POST is not. The wrapper above would happily retry a POST that timed out or got an HTTP 502 or HTTP 504, which is why the idempotency section below exists.
  • Giving up is a design decision. When retries run out, put the record on a failed list with the error, finish the rest, and report both.

Cursor pagination that survives a restart

Offset pagination (?offset=200&limit=100) breaks when records are added or removed while you page: everything shifts, so you see some records twice and skip others. A cursor is an opaque marker the server hands back that says “continue after this record”, so it doesn’t shift.

The second half of the habit is saving the cursor, so a crash halfway through resumes where it stopped, not at the start:

import json
import os

def load_cursor(path):
    if not os.path.exists(path):
        return None
    with open(path) as f:
        return json.load(f)["cursor"]

def save_cursor(path, cursor):
    tmp = path + ".tmp"
    with open(tmp, "w") as f:
        json.dump({"cursor": cursor}, f)
    # atomic: old file or new, never half
    os.replace(tmp, path)
def sync(get_page, upsert,
         path="cursor.json"):
    cursor = load_cursor(path)
    while True:
        # retries live inside get_page
        page = get_page(cursor)
        for record in page["data"]:
            # safe to see twice
            upsert(record)
        cursor = page["next_cursor"]
        if cursor is None:
            break
        # only after the page is stored
        save_cursor(path, cursor)
    if os.path.exists(path):
        os.remove(path)

What to point at:

  • The cursor is saved after the page is stored, never before. Save it first and a crash mid-page loses those records for good. Save it after and a crash repeats at most one page. That is at-least-once delivery, and it is the right trade.
  • So the write must be an upsert. Because a page can arrive twice, upsert writes by the record’s id, and a repeat overwrites instead of duplicating.
  • Retries live inside get_page. The pager never sees a transient error. One try around the whole loop, restarting from the first page, is the mistake to avoid.

The test: ten records, three per page, and a fake that dies on its second call.

records = [{"id": i} for i in range(10)]
fetched, writes = [], []

def get_page(cursor, die_on=None):
    start = int(cursor or 0)
    fetched.append(str(start))
    if len(fetched) == die_on:
        raise RuntimeError("died")
    end = start + 3
    nxt = str(end) if end < 10 else None
    return {"data": records[start:end],
            "next_cursor": nxt}

try:
    sync(lambda c: get_page(c, die_on=2),
         writes.append)
except RuntimeError as e:
    print(e, open("cursor.json").read())
# died {"cursor": "3"}
fetched.clear()
sync(get_page, writes.append)
print(fetched, len(writes))
# ['3', '6', '9'] 10

The rerun starts at the saved cursor, never refetches the first page, and writes every record exactly once.

One follow-up to expect: “What if the cursor expires?” Say it depends on the order. If the API pages in an order you can filter on, such as creation time, resume from the newest record you stored (created_after). If not, rescan from the first page, which the upsert makes safe. A max updated_at watermark only works when the pages come back sorted by updated_at. The cursor pagination question goes further into this.

Idempotency keys and the retried POST

Here is the classic mistake. The client sends a POST to create a permit. The server writes the row, then the response is lost on the way back. The client sees a timeout, retries, and the server creates a second permit. The customer now has two.

A fake API that does exactly that:

class FakePermitAPI:
    """Commits a write, loses the reply."""
    def __init__(self):
        self.permits, self.by_key = [], {}
        self.lose_reply = True

    def post(self, body, key=None):
        if key in self.by_key:
            # a replay: no new row
            return 201, self.by_key[key]
        n = len(self.permits) + 1
        permit = {"id": n, **body}
        self.permits.append(permit)
        if key:
            self.by_key[key] = permit
        if self.lose_reply:
            self.lose_reply = False
            # written, but the reply is lost
            raise TimeoutError
        return 201, permit

And the client:

import uuid

def create_permit(api, body,
                  use_key=True, attempts=3):
    key = None
    if use_key:
        # made once, before any attempt
        key = str(uuid.uuid4())
    for _ in range(attempts):
        try:
            return api.post(body, key=key)
        except TimeoutError:
            continue
    raise RuntimeError("gave up")

Run it with use_key=False and len(api.permits) is 2. Run it with the key and it is 1, and the retry gets back the permit the first attempt created. An idempotency key is a unique value the client sends with a write so the server can recognize a repeat and return the first result instead of acting again. The header has a name to cite: the IETF HTTP APIs working group drafted it as Idempotency-Key in an Internet-Draft that has since expired, and APIs such as Stripe’s use the same idea.

The subtle version of the mistake: minting the key inside the loop. Move uuid.uuid4() inside the for loop and every attempt carries a fresh key that the server treats as a new request. Then len(api.permits) is 2, as with no key. One logical operation gets one key, made before the first attempt.

Better still, derive the key from the business event rather than a random value, so it survives your own process restarting. The vercel-solutions contract above uses the event id as the idempotency key. Source 4HTTP contractPublisherVercel (vercel-solutions on GitHub)Source typecompany website For a spreadsheet sync, the row id plus the row’s last-edited value works the same way.

If the API accepts no key, offer the fallback: before retrying a timed-out create, look the record up by a natural key, such as the address and date, and create it only if it is missing. Say its limits out loud: it is only safe with one writer, and only if the lookup is guaranteed to see a write that just landed. A search endpoint that lags will report the permit missing, and you create it twice. If either is in doubt, mark the record unknown and have a person reconcile it rather than creating it blind. The receiving side has the mirror-image problem; the duplicate webhook question, also free, covers it.

How to narrate it so the interviewer hears judgment

What matters in the room is whether the interviewer hears why each line is there. This is how we teach it, not a script any employer publishes.

Open with the layers. “I’ll build this in three layers: a transport that makes one request with a timeout, a retry policy around it, and a pager on top that never sees a retry. Writes get an idempotency key.”

Name each decision as you make it. “I’m not retrying HTTP 400; the same request will fail the same way.” “The key is made outside the loop, because a new key per attempt is the same as no key.”

Common mistakes and the fix:

MistakeFix
Retrying every errorRetry timing failures only
A fixed sleep(1) between triesBackoff with full jitter
Retrying foreverA retry limit, then a failed list
One try around the whole pagerRetries per request, cursor saved
A new key on every attemptOne key per operation
Testing against the real clockInject sleep and the transport

Close with what you’d do next. In the last minutes, list what you left out: a rate limiter, date-form Retry-After, concurrency limits, retry metrics. The Coframe take-home asks for edge cases and productionization in the recorded video. Source 2Coframe Forward Deploy — Take-HomePublisherCoframe (GitHub)Source typecompany website

If the round goes sideways because requests fail and you can’t see why, the post on debugging failing API requests covers the diagnosis. If yours is a take-home rather than a live round, the FDE take-home post covers scoping, the README and the video.

Practice it before it counts

Reading this is not typing it with someone watching. Set a timer and build the retry wrapper, the resumable pager and the idempotent create against your own fake API. Then take the full prompt in the flaky API integration question, which adds follow-ups on page numbers, restarts and conflicting rate limits, and the safe retry client question, which asks you to put it all into one client. Both are in Pro. Our bank holds 180 interview questions, each with a model answer. The 7-day free trial covers a week of practice. Once the code is automatic, practice the other half of the job in the free practice case: a customer who needs something built and hasn’t told you everything yet.

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineerGlossaryExponential backoff with jitterRetrying with growing, randomized delays so that clients recovering from the same failure do not retry in lockstep.More on Exponential backoff with jitterGlossaryIdempotency keyA client-supplied identifier that lets a server apply a repeated request once, making retries safe.More on Idempotency keyGlossaryForward deployed software engineerPalantir’s title for its FDE role, called Delta internally; OpenAI and EY also post FDSE titles, each with its own duties.More on Forward deployed software engineerGlossaryBackfillReprocessing historical data through a pipeline, ideally with the same idempotent code as the daily run.More on Backfill

Questions people ask

What is an API integration coding interview?

A coding round or take-home where you connect to an API you do not control, often one that paginates, rate-limits or fails at random, and build something reliable on top of it. Public FDE take-home repositories include a mock permit API built with intentional issues and a two-way integration between a Google Sheet and a mock approval API.Source 1Permit System APIPublisherForerunner Industries (GitHub organization)Source typecompany websiteSource 2Coframe Forward Deploy — Take-HomePublisherCoframe (GitHub)Source typecompany website

Which errors should a retry wrapper retry?

Retry timeouts, connection errors, rate-limit responses and temporary server errors, and honor a Retry-After header when the server sends one. Do not retry most client errors, such as a malformed request or a forbidden resource, because the same request will fail the same way. An expired token is the exception: refresh it once, then retry once.

Is it safe to retry a POST request?

Only if the server can recognize the retry. Send an idempotency key so the server returns the first result instead of creating a second record. Without one, a timeout after the server has committed the write turns your retry into a duplicate.

Why add jitter to exponential backoff?

Without jitter, every client that failed at the same moment retries at the same moment, and a recovering service is hit by synchronized waves of traffic. Random jitter spreads the retries out.

Keep reading