Where it comes from

How to answer

Sierra’s engineering blog says it is piloting a debugging interview in which candidates improve a colleague’s draft pull request in a medium-sized codebase, using coding agents. Source 3The AI-native interviewPublisherSierraSource typecompany blog An anonymous March 2026 write-up on gaijineer.co, a blog that sells interview prep and compiles anonymous stories, describes a Sierra debugging round with a small codebase plus a diagram of how the agent should work, which was “your source of truth”, and “about 3 bugs to find”. Source 4Sierra Software Engineer, Agent Interview ExperiencePublishergaijineer.co (a blog that promotes the furustack interview-prep product)Source typeprep-product blog (anonymous, may be compiled) Neither role is titled , but the skill is the same one you use when a customer’s agent does what its design forbids.

  1. Turn the diagram into a table before you read the code. One row per edge: from, condition, to, side effect. Include edges drawn from “any state”. The table catches a missing branch, which reading code cannot.
  2. Find where each decision lives. Name the state variable and the dispatcher, then split what the model decides (the user’s intent) from what code decides (guards, limits, tool calls).
  3. Check both directions. Every edge against the code, then every branch against the diagram. Look for missing or extra edges, wrong targets, wrong guards (> for >=, or for and), misplaced side effects and checks in the wrong order.
  4. Stub the model. A scripted classifier makes transitions deterministic, so a failure points at code, not sampling.
  5. Write one test per edge, then fix one bug at a time. Rerun the whole table after each fix.
  6. Report, and ask where the diagram is silent. Give each bug’s edge, line and fix. Where the diagram is silent, pick the option with no side effect and say so.

Finish the table even after you reach the count you were given. The trap is fixing what looks wrong line by line, and trusting a rule in the system prompt as if code enforced it.

GlossaryAgentA system in which a model chooses steps and tool calls to complete a task, within limits the design sets.More on AgentGlossaryForward deployed engineerA software engineer who builds and ships production systems inside a customer’s problem and environment, accountable to that customer’s outcome.More on Forward deployed engineer

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • You have found as many bugs as the interviewer said to expect. Are you done?
  • The system prompt describes the refund policy too. Why isn’t that enough?
  • One transition depends on the model reading the user’s intent. How do you test it without calling the model?
  • Which of these bugs would a replay of recorded conversations have caught, and which only a test per edge?

Where answers go wrong

  • Reading the code top to bottom and fixing whatever looks odd. A missing transition has no line to look odd, so only a check against the diagram finds it.
  • Stopping at the number of bugs you were told to expect, when the diagram, not the count, is the spec.
  • Treating a rule written in the system prompt as enforced. If the diagram draws a guard, the code has to hold it.
  • Fixing everything in one edit with no tests, so you cannot show which change repaired which edge.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Model answer

The codebase is a support agent for a subscription product. prompt.py, llm.py and tools.py hold the classifier prompt, the model client and the tool wrappers; flow.py is where the diagram should live. “Before I open the code, I’ll write the diagram down as edges, so I have something to check against.”