A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
Bad model output must never reach your caller. Say so early: “The caller gets an object that passed every check, or a typed failure with the raw text and the reason. Never something in between.”
- Name the failure layers. Not JSON, JSON that fails the schema (a missing field, a string for a number), and a valid object that breaks a business rule (line items that don’t sum to the total, a date after today).
- Extract narrowly. Check the stop reason first: output cut off at the token limit, or a refusal, is not fixed by the same call again. Strip a code fence and parse with
json.loads. Structured output mode helps, but still validate. - Collect every schema error with its path. Pydantic reports
items.2.qtywithInput should be a valid integer(Pydantic validation errors). The path makes the retry useful. - Retry with the evidence, where it helps. Append the bad output as the assistant turn and the errors as the next user turn: “That failed validation: ... Return only the corrected JSON.” Cap attempts and tokens. Retry syntax and schema errors only. A business-rule failure goes back to the caller flagged, because telling the model “the date can’t be in the future” invites it to invent one that passes.
- Test with a scripted fake model. Bad then good returns after two calls; bad every time fails after exactly the cap; the second prompt contains the error text.
Two traps. An unbounded retry of the same prompt, which at low temperature returns the same bad output until the budget is gone. And a regex “repair” that drops or coerces what it couldn’t read, so a wrong answer reaches the caller looking valid. The model JSON repair drill runs this against a timer.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- The retry fixes the field you flagged and changes a field that was right the first time. How would you catch that?
- The business-rule check fails on one field out of many. Do you return the rest, and how does the caller know which values to trust?
- In production, how would you know how often the first attempt fails, and what would you do with that number?
Where answers go wrong
- Retries in an unbounded loop, or resends the same prompt without the error, so a model at low temperature returns the same bad output until the budget runs out.
- Repairs the text with regexes or eval until something parses, so a wrong answer reaches the caller looking valid.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Model answer
“I’ll pin the contract first. The caller gets one of three results: ok with an object that passed the schema and the business rules, flagged with a valid object and the rule it broke, or failed with every raw output and the last errors.