Where it comes from

How to answer

This question checks one thing: whether your “on purpose” survives a look at your README and commit history. A take-home review is where your decisions get examined one at a time. Sierra’s blog on its engineering onsite says that after the candidate demos what they built, the interviewers debate the key product flows and choices the candidate made, and review the code to understand their technical judgment. Source 3The AI-native interviewPublisherSierraSource typecompany blog So separate two things before you answer: what you chose not to build, and what you didn’t notice.

Prepare two lists before the round. Reread your submission as a reviewer would, run it against inputs you didn’t test, and sort what you find.

  • Deliberate cuts. Each one has the reason tied to the brief, what the cut costs, and the signal that would make you add it. Ideally the README already says so; that is your evidence.
  • Misses. Things you found after submitting. Say how you found each one and how bad it is.

Don’t change the submitted code. Put your fix for the biggest miss on a separate branch you can show if asked. When the interviewer finds something you didn’t list, say whether it’s new to you, name the fix, and say where it sits in your ranking. Don’t argue it into a cut.

Then answer in this order.

  1. Cuts first, briefly: “I left out auth and a queue on purpose. The brief said a single user and a demo, and both are in the README’s limitations section.”
  2. Then the biggest miss, owned plainly: “After submitting, I found...” Name the failure, the cause and the fix.
  3. Then the next day, ranked, and say the rule you ranked by out loud, for example the cost of the failure times how likely it is, weighted up when nobody would notice it. Put the miss where the rule puts it, not automatically first.
  4. Close on production: what would still stand between this and a real customer.

Spoken, the answer below runs about three and a half minutes. Stop there and let them pick what to dig into.

The trap is relabeling a miss as a cut. A real cut was decided before the deadline and can leave a trace in the README or the commit history, so a relabeled miss can be checked, and if one claim fails, the others are in doubt too.

Follow-ups

What the interviewer may ask next, once your first answer is on the table.

  • You say you left authentication out on purpose. Where does the README say that?
  • Why is that first on your list and not the second item?
  • Which of those changes would you make before a real customer ran this?

Where answers go wrong

  • Presenting something you missed as a deliberate cut. The commit history and the README show which one it was.
  • Listing improvements as an unranked wish list of tests, Docker and a nicer UI, with no reason for the order.

Answer this in two minutes

Write the answer you would say out loud. The clock starts with your first word.

Two minutes

Compare with the model answer

Illustrative answer about a fictional project

Written in the first person to show the structure. Tell yours from your own work.

“The brief was a ticket triage service: classify incoming support tickets into one of eight queues and draft a first reply with an LLM, for a support lead who reviews the drafts. I’ll give you what I left out on purpose, the biggest thing I missed, and then how I’d spend another day.

“Left out on purpose. Three things, all in the README’s limitations section: retries on the model API, authentication and a background queue. The queue is the one worth explaining. The brief said about 300 tickets a day, so classification runs inline with a timeout=20s, and a timeout returns the ticket unrouted with a flag rather than guessing.

“The biggest miss. After submitting, I ran it on the 200 unlabeled tickets in the sample file and read the output. Eleven tickets in Spanish or Portuguese went to the ‘other’ queue with a self-reported confidence above 0.9. I’d used the model’s own score as the floor, and it isn’t calibrated: my eval set had no non-English tickets, so nothing checked it, and the floor only catches answers the model itself rates as unsure. A confident wrong route is the worst failure for the support lead, because nobody looks at it.

“Another day, ranked by what a failure costs the support lead, how likely it is, and whether anyone would notice.

  1. Add a multilingual slice to the eval set, and route anything outside the languages we’ve tested to a human queue. Then replace the self-reported score. Where the model API exposes logprobs, I’d use the logprob of the label’s first token, with queue names chosen to start with distinct tokens; where it doesn’t, how often five samples at a temperature above zero agree. Either way, the threshold is tuned on the eval set. It’s the only failure I’ve seen that is both likely and silent.
  2. Make ingestion idempotent on the ticket ID. I found this in the same post-submission pass, so it’s a miss: if the helpdesk webhook retries, today we draft two replies. That’s likely with any real webhook, but the support lead would see the duplicate, so it’s second.
  3. Retry HTTP 429 and 5xx responses from the model API with exponential backoff and , capped at max_retries: 4. This is the retry cut from the README: right now a rate limit surfaces as an unrouted ticket, which is safe but annoying, and the support lead sees it, so it’s third.

“What I’d still leave out after that day: the UI and a second model for cost routing. Neither changes whether the support lead can trust the queue.

“Before a real customer ran it, I’d want auth tied to their helpdesk accounts, and a real eval set. Mine is 60 tickets from the sample data that I labeled myself, which shows the pipeline works, not how accurate it would be in production; theirs would come from their last month of tickets, labeled by their team. And a weekly report of how often the lead edits the drafts, so we’d know whether it’s saving them time.”

If they ask where the README says it: “Limitations, the second bullet: ‘No auth: single local user per the brief.’ It’s in the last commit before the deadline, and I can show you.”

If a cut you name isn’t written down: “It isn’t in the README, so treat it as a miss. Here’s where it ranks.”

GlossaryExponential backoff with jitterRetrying with growing, randomized delays so that clients recovering from the same failure do not retry in lockstep.More on Exponential backoff with jitter