In this post12 sections
  1. The scenario, and what it tests
  2. Scope the blast radius before you debug anything
  3. Build a timeline and check what changed
  4. Read the errors by class: client, server, throttled or unavailable
  5. Restore service first, find the root cause second
  6. The status update you send the customer
  7. After the hour: root cause, follow-ups and the write-up
  8. Mistakes that end the round early
  9. Practice the first hour out loud
  10. Questions people ask
  11. Keep reading
  12. More from the blog

The interviewer leans back and says: “A customer just called. Their requests to our API are failing, their operations team is escalating, and you have the dashboards. What do you do?” There is no code on the screen yet, a clock is running, and you can feel the pull to open the first stack trace and start guessing. The way to debug failing API requests in an interview is to work in a fixed order, out loud: scope the blast radius, build a timeline, check what changed, read the errors by class, restore service before you chase the root cause, and send the customer an update with a time for the next one. This post walks that first hour step by step, with the words to say at each one. For the rest of the loop, round by round, start with the FDE interview guide.

The scenario, and what it tests

Here is the scenario in a form an employer has published. Ethyca’s take-home for its Forward Deployed Privacy Engineer role has you connect to its API, reproduce and diagnose a customer’s API error, and then write the email reply to that customer. Source 1Ethyca Technical Challenge -- Forward Deployed Privacy Engineer (FDPE)PublisherEthyca (GitHub)Source typecompany website You may instead get logs, an architecture diagram or only a spoken prompt; the order below works on all three.

Employers say they care about the skill underneath. Palantir’s page on working inside existing systems says “a developer will spend far more time debugging, modifying, and extending existing code than writing brand-new things from the ground up.” Source 2Working Inside Existing SystemsPublisherPalantirSource typecompany hiring page Ramp’s engineering blog says its hiring validates fundamentals including coding, debugging, data structures and system design. Source 3Forward Deployed EngineeringPublisherRamp Builders (engineering blog)Source typecompany blog The job asks for it too: OpenAI’s financial-services FDE posting asks the FDE to own reliability, observability and on-call readiness Source 4Forward Deployed Engineer (FDE), Financial Services- NYCPublisherOpenAI (Ashby)Source typecompany job posting, and Okta’s Senior FDE posting asks for on-call experience. Source 5Senior Forward Deployed Engineer - Okta for AI AgentsPublisherOkta (Greenhouse)Source typecompany job posting

Debugging can also be a round of its own. A candidate report for a Scale AI FDE onsite, dated August 2025, lists “debugging a giant repo” as one of five parts. Source 6Scale AI Forward Deployed Engineer Interview Experience (2026)PublisherAced (formerly Exponent)Source typecandidate’s personal write-up One poster wrote on Reddit in August 2026 to ask whether anyone had been through a Cohere FDE “AI coding and system debugging round”. Source 7Cohere FDE interview AI coding and system debugging round - has anyone been for it before and is able to share their experience? thank you! (post by u/Exciting-Art6805)PublisherReddit r/leetcodeSource typecandidate report on Reddit Sierra says it is piloting a debugging interview for its engineering hiring, not specifically FDE, in which candidates improve a colleague’s draft pull request in a medium-sized codebase using coding agents. Source 8The AI-native interviewPublisherSierraSource typecompany blog In a codebase, the order below shrinks to four moves: reproduce the failure, count where it happens, form one hypothesis, and test one change.

No employer we have read publishes a rubric for this scenario. The first-hour order below is our method, not a checklist an interviewer holds. It gives you a spine, so every question the interviewer throws at you lands somewhere in the plan instead of knocking you off it.

Open by saying the plan in one breath:

“I’ll work in this order: find out who is affected and how badly, line up when it started against what changed, split the errors by type, get the customer working again, and only then dig for the root cause. I’ll keep the customer updated the whole way.”

That sentence alone shows you won’t spend the hour inside one stack trace.

Scope the blast radius before you debug anything

“Requests are failing” is a feeling, not a measurement. Your first job is to turn it into numbers you can query. Ask these before you touch a log line:

Scope questions for the first minutes

  • Who? One customer, one region, one plan tier, or everyone?
  • What? Every endpoint, or one route, one method, one SDK version?
  • Since when? The time of the first failure, with the time zone, not the time of the first complaint.
  • How much? The error rate, meaning failures divided by all requests, not a raw count.
  • How bad? Hard errors, slow responses, or wrong answers returned with success codes? Is any data lost or written twice?
  • Evidence? A few failing request IDs and the exact error body the customer’s client saw.

A jump in errors can be nothing more than a jump in traffic, so use the rate. Then pull the same rate for everyone else, and the scenario forks:

  • Only this customer fails. Look at what they send and how their account is set up: a rotated key, a new payload field, a quota they hit, a new client release.
  • Everyone fails on one route or one host. It is yours, and it is an incident. Say you would page whoever is on call before going deeper.

One trap is worth naming aloud. A request that never reached your service leaves no error in your logs. If the customer’s request IDs are missing, check the load balancer, the gateway and the TLS handshake. If they appear with a success code and a long duration, the customer’s client gave up before you answered, and an error dashboard will never show it.

Say it like this: “Before I debug, I want the error rate for this customer over time and the same number for everyone else. That tells me whether this is their integration or our outage, and how loud the response needs to be.”

Build a timeline and check what changed

A recent change is the first suspect in any outage. Build a timeline, even a rough one, and write it where the interviewer can see it:

07:15  deploy v2.41 (canary 10%)
09:15  promoted to 100%
09:40  first 503s, POST /v1/orders
09:55  error rate climbs, p99 doubles
10:00  customer calls

Then walk the list of things that change, in roughly this order:

  1. Your deploys, including the ones behind a canary or a feature flag.
  2. Configuration: flags, timeouts, pool sizes, rate-limit settings, routing rules.
  3. Secrets and certificates: a rotated API key, an expired TLS certificate, a revoked token.
  4. Quotas and limits: a plan change on the customer’s account, or a new limit on a provider you depend on.
  5. Dependencies: a database failover, a model provider’s slowdown, a DNS change.
  6. Traffic: a batch job, a product launch, a retry loop in someone’s client.
  7. The customer’s side: a new SDK version, a new egress IP, a changed payload.

The last item is the easiest to forget. An FDE sits between two systems, and the change that broke things can be on the other side. Ask the customer, politely and early, “Did anything change on your side in the last day: a release, a new job, a key rotation?”

When a change lines up with the first error, say so and hold it as a hypothesis, not a verdict: “The promotion to full traffic lines up with the first errors. That is my lead suspect. Before I blame it, I want to see whether the errors are only on the new version.”

Worked example: one customer, nothing changed on our side

Now the interviewer pushes back. Here is how the order survives it:

Interviewer: “Only Acme fails, and we shipped nothing.”

You: “Then I diff Acme’s failing requests against their successful ones: status, route, body size, SDK version, source IP.”

Interviewer: “Every failure is HTTP 413 on POST /v1/orders, starting at 09:40.”

You: “Then their payloads grew. I’ll send them the size limit and a workaround, splitting the batch, right now, and ask what they changed at 09:40.”

Each answer is one query and one conclusion. You never guessed; you narrowed.

Read the errors by class: client, server, throttled or unavailable

A status code tells you which side to look at. The status code classes in RFC 9110 split 4xx client errors from 5xx server errors, and that split is your first slice.

CodeWhat it saysFirst move
400, 422The request is wrongDiff failing payloads against good ones
401Missing or bad credentialsCheck key rotation
403Valid key, not allowedCheck scopes and plan
413Body too largeCompare request sizes
429Rate limitedFind which layer sent it
500Unhandled errorRead the stack traces
502, 504A proxy got a bad or late answerCheck the upstream and one host
503Server can’t serve nowCheck load, health and maintenance

If you sit in front of a model API, add three suspects. The provider’s own rate limit shows up as HTTP 429 from the layer behind you. Requests that grew past the context window fail as an HTTP 400 that started when the customer’s prompts got longer. And long generations can outlive your gateway timeout, giving an HTTP 504 on a request the model is still answering.

The pair to know cold is HTTP 429 against HTTP 503, because they call for opposite fixes.

HTTP 429 Too Many Requests means the client sent too many requests in a given amount of time, and the server may send a Retry-After header saying how long to wait (RFC 6585, section 4). The client should slow down. Your questions: which layer sent it (your gateway, your app, a provider behind you), what the limit is, and whether this customer’s traffic changed.

HTTP 503 Service Unavailable means the server can’t handle the request right now, because of a temporary overload or scheduled maintenance, and it may also send Retry-After (RFC 9110, section 15.6.4). The fault is on your side. Your questions: are instances failing health checks, is a dependency down, is a pool exhausted?

Both lead to the same danger. Clients that retry fast and in step can keep a recovering service down, which is why exponential backoff with jitter and circuit breakers exist. If a service came back and fell over again, you are looking at the pattern in the retry storm question.

HTTP 504 on a write needs a warning of its own. RFC 9110 defines it as the gateway not getting a timely answer from upstream (section 15.6.5); it says nothing about whether the upstream finished the work. And POST is not idempotent (section 9.2.2), so some of those orders may exist. If the customer retries blindly, you can end up with duplicates, the problem behind the duplicate webhook question. The coding side of that, retries and idempotency keys, is in the API integration coding round post.

If you have raw logs, bucket before you read. A tiny function is enough:

from collections import Counter

def error_class(status: int) -> str:
    if status == 429:
        return "throttled"
    if status in (502, 503, 504):
        return "unavailable"
    if 400 <= status < 500:
        return "client"
    if status >= 500:
        return "server"
    return "ok"

statuses = [200, 503, 200, 429, 401,
            200, 504, 503, 500, 200]
counts = Counter(map(error_class, statuses))
for cls, n in counts.most_common():
    print(cls, n, f"{n/len(statuses):.0%}")

On that sample it prints ok 4 40% first, then unavailable 3 30%, then one each of the rest. Then slice the largest class by route, host, client version and time. The point is not the code; it is that you counted before you read.

Restore service first, find the root cause second

Under pressure, this is the step people skip. The customer does not need the root cause in the first hour. They need their requests to work.

Say the rule plainly: “If a recent change lines up with the start of the errors and rolling it back is safe, I roll it back now, and I find out why afterward.”

Your mitigation options, from least to most drastic:

  • Turn off a feature flag that shipped the new path.
  • Roll back the deploy, or stop the canary release before it goes further.
  • Drain one bad host from the load balancer.
  • Raise a limit or shed load for a traffic spike.
  • Fail over to another region or a standby dependency.
  • Give the customer a workaround: smaller batches, a slower send rate.

Then name when a rollback is not safe; it is the obvious follow-up. A deploy that ran a database migration may not roll back cleanly. And if the errors started before the deploy, rolling it back fixes nothing and costs you time.

Before you roll back, keep the evidence: the logs for the window, the bad build and a sample of failing requests.

The status update you send the customer

The interviewer may play the customer here, or simply ask, “What do you tell them?” Either way, have the shape ready; Ethyca’s take-home ends with exactly this, an email to the customer. Source 1Ethyca Technical Challenge -- Forward Deployed Privacy Engineer (FDPE)PublisherEthyca (GitHub)Source typecompany website Write it the way you would send it, continuing the timeline above:

Since 09:40 UTC, part of your POST /v1/orders traffic has failed with HTTP 503 and HTTP 504 errors. The cause is on our side: a release that reached all traffic at 09:15 UTC. We rolled it back at 10:20 UTC and errors stopped at 10:22 UTC.

Requests that got an HTTP 504 may still have created orders. Before retrying, check the attached list, or send an Idempotency-Key header so a retry cannot create a duplicate.

We are finding out why the release failed. Next update by 14:00 UTC today, sooner if anything changes.

The twenty minutes between the call and the rollback are the hour playing out: scope, timeline, error classes, then the decision.

Every update carries four things:

  1. Impact in the customer’s terms: which of their calls, since when.
  2. What you know and what you don’t. Never state a root cause you haven’t confirmed.
  3. What they should do, especially whether retrying is safe.
  4. When they’ll hear from you next, as a time, and then keep it.

Send the first update early, even if it only says “investigating, next update in thirty minutes”. Silence turns an incident into an escalation. The lessons on writing for customers and delivering bad news go deeper. When the customer is angry rather than worried, read how to answer the customer escalation question.

After the hour: root cause, follow-ups and the write-up

Once service is back, the interviewer may ask, “And then?” This is where you show you close the loop.

  • Find the root cause with the evidence you kept. Ask why the change broke things, and then why nothing caught it before the customer did.
  • List follow-ups with owners. An alert on the error rate against a service level objective, so you hear first next time. A test for the failure. Backoff in the client, if a retry loop made it worse.
  • Write it up without blame: timeline, impact, what went well, what didn’t, follow-ups. Send the customer a version too.

If the interviewer turns it into a behavioral question, “tell me about an incident you led”, the same order works as the structure of your story. That question has the framework.

Mistakes that end the round early

  • Debugging the first stack trace for the whole hour. Fix: count the errors by class first. That trace may be background noise that was always there.
  • Blaming the customer too early. Fix: check your own changes and your own logs before you ask what they changed, and ask without accusing.
  • Going silent on the customer. Fix: give a next-update time in your first reply.
  • Telling the customer to “just retry”. Fix: say which calls are safe to retry, and how to avoid duplicates on writes.

Practice the first hour out loud

Reading the order is not saying it while someone interrupts. Take the failing requests question and talk through it with a timer, then try a follow-up: the request IDs aren’t in your logs. That question is free, with a model answer that walks the hour in timed blocks, with the jq queries to run and the update to send. When you can say the order without notes, run the free practice case, where a customer answers back and only tells you what you ask (free with a sign-in).

GlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineer

Questions people ask

What should you do first when a customer reports failing API requests?

Scope it before you debug it: which customers, endpoints and regions, since when, and what share of requests fail. Then build a timeline and check what changed, including deploys, configuration, certificates, quotas and the customer’s own changes.

What is the difference between HTTP 429 and HTTP 503?

HTTP 429 Too Many Requests means the client is being rate limited, so it should slow down and honor Retry-After if the server sends it. HTTP 503 Service Unavailable means the server cannot handle the request right now, for example because it is overloaded or down for maintenance.

Should you roll back before finding the root cause?

If a recent change lines up with the start of the errors and rolling it back is safe, yes. Restoring service comes first; keep the logs and the bad version so you can find the root cause afterward.

Keep reading