A practice prompt we wrote. No company or candidate report names it, so it carries no company tag.
How to answer
An is a promise about what the customer’s users experience, so start from their journey, not from the model.
- Ask what the assistant is for and who notices when it fails. “Claims handlers draft replies with it inside their ticket tool” gives you the moments that matter: waiting for the draft, getting none, and getting a wrong one.
- Define three indicators as good events over valid events. Availability: requests that return a usable answer, not an error, a timeout or an empty response. Latency: time to first token if the interface streams, and total time. Quality: the share of sampled production responses that pass a written rubric. Say out loud what counts as valid. A cancel in the first second or two is noise; a cancel after the latency threshold is a bad event, because slowness is why people cancel.
- Explain how quality gets measured. A daily sample of production traffic, graded against the rubric by people or by a you check against people. Add cheap signals on every request: whether cited sources exist in what was retrieved, and whether the user sent the answer or rewrote it.
- Set targets from a baseline, then attach an error budget and a policy. Measure before promising. When a budget is spent, prompt and model changes stop except fixes. Alert on burn rate, as the SRE Workbook chapter on alerting on SLOs describes. Keep harms that must never happen out of any budget: each one is an incident.
- Own the dependency. A provider outage counts against the customer’s SLO, because their users don’t care whose fault it was. That argues for a fallback model.
The trap is the endpoint uptime number. Say early that a fast, successful, wrong answer is the failure you most need to catch.
Quality sampling, model graders and drift are taught in Evaluation: proving it works, and change control, observability and rollback in a customer’s environment in Enterprise system design for FDEs. That module’s first lesson, enterprise design is different, is free; the rest of both modules is in Pro.
Follow-ups
What the interviewer may ask next, once your first answer is on the table.
- Your model provider has an outage. Does it count against your SLO?
- How do you know the grader behind your quality number is still right?
- The quality budget is spent halfway through the month. What changes the next morning?
Where answers go wrong
- An uptime target on the API endpoint and nothing else. The assistant can return fast, successful, wrong answers all week while every dashboard stays green.
- Measuring quality only on an offline eval set at release time, so a drift in production inputs never shows up in the SLO.
Answer this in two minutes
Write the answer you would say out loud. The clock starts with your first word.
Compare with the model answer
Model answer
“I’ll take a concrete case: an assistant that drafts replies for an insurer’s claims handlers from policy documents, in their ticket tool, which streams the draft as it’s written. The handler waits for the draft, reads it and sends it or rewrites it. Three things hurt them: no draft, a slow draft and a wrong one. So three indicators, each defined as good events divided by the events that should count, the ratio the SRE Workbook’s chapter on implementing SLOs recommends.”
“Concretely: 99.5% of requests return a usable draft over 28 days, 95% show the first words within 2s, and 90% of sampled drafts pass the rubric over a rolling week. Harm has no budget. Written as a spec:”
slos:
availability:
good: >-
draft returned, non-empty,
no error, within 30s
valid: >-
all requests except malformed
and user-canceled within 2s
target: 99.5%
window: 28d rolling
latency:
time_to_first_token:
threshold: 2s
target: 95%
full_draft:
threshold: 15s
target: 99%
bad: >-
canceled after 15s counts
against latency
window: 28d rolling
quality:
sample: >-
200 production drafts per day,
stratified by claim type
good: >-
passes rubric (grounded in
cited clauses, answers the
question)
target: 90%
window: 7d rolling
harm:
examples:
- coverage promised outside policy
- another claimant's data
target: zero
response: >-
incident and postmortem,
not budget
online_signals:
citations:
good: >-
every clause the draft cites
exists in the retrieved passages
measured: every request
acceptance:
good: >-
draft sent with light edits
(edit distance under a threshold)
measured: every request
“Time to first token is there only because the tool streams. If the draft dropped in whole, full-draft time is the only latency the handler feels, and I’d drop that indicator. The targets come from a baseline I’d measure for a couple of weeks first, and I’d agree them with the customer’s operations lead, not pick them myself. The quality window is 7 days because that’s about 1,400 graded drafts, enough that a drop of a few points is signal rather than noise. A daily number at 200 moves that much by chance. The window is still short enough to catch a prompt change within the week.”
“Quality needs the most care. The rubric is written down, with examples of pass and fail. A grades the daily sample, and each week two of the customer’s senior handlers grade 50 of the same drafts, which add up to a monthly agreement figure between them and the judge. If agreement drops, the quality number is suspect before it’s low, and I recalibrate the judge before trusting it again. Every draft that fails becomes a case in the that gates the next change.”
“The cheapest quality signal is what handlers do with the draft. A rising rewrite rate is the users telling me quality dropped, on every request, before the grader does. It can’t be the on its own, because a busy handler sends a plausible wrong draft unedited, which is why the graded sample stays. The citation check is cheap too: a draft that cites a clause the retrieval step never returned is a hallucination I can catch without a grader.”
“Harm is kept apart. Promising coverage the policy doesn’t give isn’t one draft in ten the budget tolerates; one case is an incident with a postmortem.”
“Error budgets carry a policy, or they’re decoration. When the quality budget is spent, prompt and model changes freeze except fixes, and the next change needs a full regression run instead of the fast subset. The morning the quality budget runs out, I pull yesterday’s failed drafts, group them by failure type, and name the top one and who owns the fix in the customer’s standup. The freeze goes out in the project channel that morning, with what lifts it: the rolling week back above target. For availability I’d use the Workbook’s multiwindow, multi-burn-rate alerts: page at 14.4x burn over 1h (confirmed over 5m), which spends 2% of the budget in an hour, or at 6x over 6h (confirmed over 30m), and open a ticket at 1x over 3d (confirmed over 6h). The Workbook derives those rates for a 30d window, so I’d rescale them for our 28d one.”
“These targets are our internal SLOs. What goes in the contract is looser, so we see trouble before the contract is breached, and I’d put quality in as a reported metric with a review process, not a service credit.”
“The model provider’s outages count against us. The customer bought an assistant, not a provider. So there’s a fallback model behind a router with its own quality sample, and if the fallback scores lower on the rubric, the customer knows that in advance.”
“What I’d leave out: model-level metrics such as token throughput. They’re useful for capacity planning, but no claims handler experiences them.”